Skip to content

Reproduction  #1

Description

@dorsa-zeinali

Hi,
I tried reproducing your code for the 2 bit quantization of llama-2-7b , and it took 4 hours and 12 minutes with 2 RTX 3090 gpus. Is this expected?. Also, for wiki text2 ppl I get 14.59, which is close to your reported value of 16.73. However, for c4, I get a ppl of 56.73, which is not close at all to your reported ppl of 23.73. Is this expected? Could there be a bug on my end? I did not get any errors during the process. Thank you and please let me know.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions