Skip to content

Any upcoming update to your repository? #1

Description

@ashim95

Hi,

Thanks for releasing the code.

Can you please answer a few questions?

  1. Does this code reproduce results from your EMNLP paper? And when do you plan to release all the preprocessing scripts along with readme?

  2. I tried preprocessing the reddit l2 dataset the way you described in the paper (posts greater than 50 words and balanced dataset for all classes). But I did not get the numbers of the dataset you reported (260k for training, 32k for test, valid). Are these numbers for 10 most frequent classes or all 23 of them?

  3. In your arxiv version, the numbers for a linear classifier (LR) in tables 1 and 4 do not match (52.5 vs 21.1 for in-domain). I assume it is because the num_classes in former is 10 while it is 23 in the latter?

  4. Can you share the exact splits (train, test, oodtest) you used in your paper?

Thanks again for releasing the code,
Ashim

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions