Hi,
I tried reproducing your code for the 2 bit quantization of llama-2-7b , and it took 4 hours and 12 minutes with 2 RTX 3090 gpus. Is this expected?. Also, for wiki text2 ppl I get 14.59, which is close to your reported value of 16.73. However, for c4, I get a ppl of 56.73, which is not close at all to your reported ppl of 23.73. Is this expected? Could there be a bug on my end? I did not get any errors during the process. Thank you and please let me know.
Hi,
I tried reproducing your code for the 2 bit quantization of llama-2-7b , and it took 4 hours and 12 minutes with 2 RTX 3090 gpus. Is this expected?. Also, for wiki text2 ppl I get 14.59, which is close to your reported value of 16.73. However, for c4, I get a ppl of 56.73, which is not close at all to your reported ppl of 23.73. Is this expected? Could there be a bug on my end? I did not get any errors during the process. Thank you and please let me know.