Skip to content

Ingest LREC and workshops - #9304

Open
mjpost wants to merge 52 commits into
masterfrom
ingest-lrec
Open

Ingest LREC and workshops#9304
mjpost wants to merge 52 commits into
masterfrom
ingest-lrec

Conversation

@mjpost

@mjpost mjpost commented Jul 25, 2026

Copy link
Copy Markdown
Member

Administrative checklist:

  • Make sure the branch is merged with the latest master branch
  • Ensure that there are editors listed in the <meta> block
  • In the Github sidebar, add the PR to the current milestone
  • In the Github sidebar, under "Development", link to the corresponding ingestion issue (if applicable)
  • Ensure ORCID iDs are present for (most) authors
  • Ensure OpenReview IDs are present for all authors if it is an OpenReview venue
  • For workshops, add a <venue>ws</venue> tag to its meta block
  • For workshops, add a backlink from the main event's <event> block
  • Add events to their relevant SIGs

Navigate to the preview site and check the following:

  • Verify that volume titles are consistent with past years (click on the venue link to see them)
  • Skim through the complete listing, looking for mis-parsed author names
  • Download the frontmatter and verify that the table of contents matches at least three randomly-selected papers
  • Download 3–5 PDFs (including the first and last one) and make sure they match the paper's metadata (title, authors, page numbers, etc)
  • Search the PDFs for "Anonymous ACL submission" or similar to identify non-final uploaded papers

After the PR is closed, for all events:

  • Archive the ingestion materials in format YYYY-MM-DD-{event}

Create followup issues for the following tasks (usually just for ACL events)

  • Create DOIs
  • Ingest videos
  • Add awards

@mjpost mjpost added this to the 2026Q3 milestone Jul 25, 2026
@mjpost mjpost self-assigned this Jul 25, 2026
@github-actions

Copy link
Copy Markdown

📊 Ingestion statistics

Analyzed 47 volume(s): 2026.bucc-1, 2026.cas-1, 2026.cawl-1, 2026.chipsal-1, 2026.cl4health-1, 2026.clinicalnlp-1, 2026.cmcl-1, 2026.cmlc-1, 2026.delite-1, 2026.determit-1, 2026.dialres-1, 2026.dmr-1, 2026.dtf-1, 2026.fnp-1, 2026.gaze4nlp-1, 2026.htres-2, 2026.iaai-1, 2026.indor-1, 2026.isa-1, 2026.kallm-1, 2026.lanlp-1, 2026.ldl-1, 2026.legal-1, 2026.llms4ssh-1, 2026.lrec-1, 2026.lt4hala-1, 2026.nakbanlp-1, 2026.neollm-1, 2026.nlp4ecology-1, 2026.nlperspectives-1, 2026.nonliteral-1, 2026.nslp-1, 2026.osact-1, 2026.parlaclarin-1, 2026.politicalnlp-1, 2026.pressmint-1, 2026.rail-1, 2026.rapid-1, 2026.readi-1, 2026.resourceful-4, 2026.signlang-1, 2026.sigul-1, 2026.slide-1, 2026.socon-1, 2026.speakable-1, 2026.udw-1, 2026.wildre-1

Metric Count
New papers 1890
Distinct authors 6023
New authors (first time in the Anthology) 2410
New people.yaml entries 647
Single-author papers by new authors 50
Paper-author instances with ORCID iD 1788 / 7703 (23.2%)

Top ten authors by paper counts: Wajdi Zaghouani (21), Georg Rehm (11), Els Lefever (9), James Pustejovsky (9), Jelke Bloem (9), Barbara Plank (8), Eiji Aramaki (7), Hafsteinn Einarsson (7), Marco Carlo Passarotti (7), Maxime Amblard (7)

Papers by number of authors

xychart-beta
    title "Papers by number of authors"
    x-axis "Number of authors" [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 21, 26, 29, 30, 78, 107]
    y-axis "Papers" 0 --> 454
    bar [156, 369, 454, 348, 213, 133, 80, 42, 29, 24, 9, 2, 9, 4, 4, 3, 1, 3, 1, 1, 2, 1, 1, 1]
Loading

@github-actions

github-actions Bot commented Jul 25, 2026

Copy link
Copy Markdown

@mjpost mjpost linked an issue Jul 25, 2026 that may be closed by this pull request
@nschneid

Copy link
Copy Markdown
Collaborator

Would it be easy to rerun with a temporary fix for the newline bug? I see ".The" in the data, for example.

@nschneid

Copy link
Copy Markdown
Collaborator

Would it be easy to rerun with a temporary fix for the newline bug? I see ".The" in the data, for example.

ingest.py is now updated on master.

@mjpost

mjpost commented Jul 28, 2026

Copy link
Copy Markdown
Member Author

ingest.py is now updated on master.

Okay, I re-ingested LREC main, and there were two small changes.

It was a bunch of work correcting mistakes, so I'd prefer to just apply this going forward from now on (and not do this for all workshops).

@mjpost

mjpost commented Jul 28, 2026

Copy link
Copy Markdown
Member Author

@nschneid Something else that has come up, is they have asked for us to point to the LREC-hosted PDFs. We still have our local copies. Ideally we could point to both but I think at the moment we have to choose one.

@nschneid

Copy link
Copy Markdown
Collaborator

It seems that starting with LREC 2016 we have been hosting PDFs. The advantage of that is more redundancy (in case one of the servers is down). The advantage of pointing to theirs would be that we would not have to worry about PDF revisions on our end. Though of course we would want to know about title/abstract/author changes or retractions.

@nschneid

Copy link
Copy Markdown
Collaborator

ingest.py is now updated on master.

Okay, I re-ingested LREC main, and there were two small changes.

Hmm, then maybe repair_latex() wasn't the (sole) source of the problem. Can you check whether ".The" is present in source data?

@mjpost

mjpost commented Jul 28, 2026

Copy link
Copy Markdown
Member Author

Hmm, then maybe repair_latex() wasn't the (sole) source of the problem. Can you check whether ".The" is present in source data?

It is. Input file cdrom/bib/2026.lrec2026-1.938.bib:

abstract = {The Romanian journalistic corpus previously annotated with verbal multiword expressions (PARSEME-Ro) has been extended recently with other journalistic texts and annotated with multiword expressions of all parts of speech closely observing version 2.0 of the PARSEME guidelines. The corpus size has been increased by about 40%, it underwent automatic morpho-syntactic annotation following the Universal Dependencies principles, as well as extensive semi-automatic annotation of multiword expressions of all morphological types (nominal, adjectival, adverbial, determiner, pronominal, prepositional, conjunction, interjection, and verbal for the newly added texts). We present here our work methodology, which involves an automatic annotation phase, but the manual work prevails in checking the annotation and its consistency. We also offer quantitative data about the new version of the corpus, the types of multiword expressions existing in Romanian and occurring therein, and characteristics thereof. The new version of the PARSEME-Ro corpus contributes to the field of developing multiword expressions resources per se, i.e. describing this language phenomenon, as well as resources for training, tuning and testing the performance of tools and large language models when dealing with this linguistic phenomenon.The paper also discusses some remarks on the MWE paraphrasing subtask in which a part of the corpus was used. The corpus is released with a permissive license.},

@nschneid

Copy link
Copy Markdown
Collaborator

OK. Based on some stats I have a hunch that the newline bug was mainly affecting aclpub2 proceedings. Maybe the LREC undersegmentations are just user typos.

@nschneid

Copy link
Copy Markdown
Collaborator

When your work on the XML is done I can run my script from #9345 as that is currently independent of the ingestion code.

Comment thread data/xml/2026.cmlc.xml
<author><first>Svetla Peneva</first><last>Koeva</last></author>
<author><first>Ivelina</first><last>Stoyanova</last></author>
<pages>12-24</pages>
<abstract>Many thanks to all reviewers for the detailed comments and suggestions. Misspelling, formatting errors and other minor issues have all been corrected, and are not listed below. 1. (Reviewer 1) Comment: For example, the relation of some described corpora to the Bulgarian National Corpus (like MIC21, Bulgarian MARCELL, General News in Bulgarian, ...) is not clear to me. Are they part of BulNC or the other way around? It’s stated that the BulNC is part of CURLICAT, so I would assume the relation with the other corpora is similar? Response: The conclusion now clarifies the relations between BulNC and IfGPT and its subsets. The paper was also restructured to clarify these issues. 2. (Reviewer 1) Comment: Some sections on the paper reference related work, where again the relation is not clear to me, like on page 3 "Another direction, which still presents significant challenges, is towards multilingual data. O’Keeffe et al. (2024) describe a pipeline for capturing professional video-call interactions, including screen recordings, speaker tracking, and facial expression data. Macaire et al. (2024) present a speech text pictogram corpus for French (230 hours), targeted at augmentative and alternative communication research. Lai and Pustejovsky (2024) develop an annotation scheme for iconic, deictic, and beat gestures anchored in Abstract Meaning Representation." As these publications follow the MIC21 corpus, it’s not clear to me, how they relate. Response: These references are reduced; a citation for the MIC corpus is provided where more details are available on the related works of MIC. 3. (Reviewer 1) Comment: I would be very interested in the benefits of a Graph database for the Metadata and I think it would be very valuable to describe, why the Graph database is more appropriate for Metadata than commonly used alternatives. The described relations are not totally convincing to me. Response: A paragraph is added in the Metadata management section on the justification of the use of a graph database. 4. (Reviewer 1) Comment: The publicly accessible web interface should be linked to. It’s also not clear to me, if it provides a fulltext search or only a keyword search in the Metadata (i.e. in the Graph database). Response: Links are provided to both the search interface of BulNC (full text search, mainly for linguistic research) and to IfGPT metadata search interface (allowing selection of subdatasets for NLP tasks and LLM fine-tuning). 5. (Reviewer 2) Comment: My first doubt is why all the data is presented as part or at least related to the Bulgarian National Corpus - to me, national corpora are reference language corpora with all the characteristics this entails, like a carefully balanced corpus representative of contemporary standard Bulgarian. I would not expect to see in this context mentioned artificially produced language for LLMs, multilingual or foreign language corpora or image corpora. To me it would make a lot more sense the re-cast the paper as presenting (newly) available language resources for Bulgarian (maybe in the context of the Bulgarian CLARIN / CLARIAH, which is not even mentioned!) rather than shoe-horning them to the BulNC. Response: The ties between the BulNC and the large dataset IfGPT has been clarified in the Introduction, in the text and in the Conclusion. 6. (Reviewer 2) Comment: Second, it would help the reader that, rather than just listing all the resources, a table with their key characteristics would be provided first, and then each introduced. Similarly, the endless repetitions of "so-and-so many JSON files" for each domain in Sec. 2.5 are not really helpful, as, first, texts or words are a better metric than files, and, second, this information would also be better presented in a table. Response: A table has now been provided listing the resources and their key properties, including size. 7. (Reviewer 2) Comment: "The BulNC-based dataset is publicly accessible through a dedicated web interface" I don’t see that these datasets are in any way BulNC-based. Response: Clarifications are made on this issue in the Introduction and the Conclusion. 8. (Reviewer 3) Comment: I am missing (at least a sketchy) description of tools used for tagging and/or parsing the corpus data. Response: A reference to the Bulgarian Language Processing Tool Set is now provided. 9. (Reviewer 3) Comment: It is also not clear why the corpus is still using the web interface developed before 2014, and not any newer tool (with a FLOSS license, such as CQPweb or NoSketch Engine). Response: This has been clarified in section 2.1. 10. (Reviewer 3) Comment: I would also expect the text to include summary statistics (preferably in a single table), so the sizes of the respective resources can be compared. Response: Table 1 is now provided for this purpose.</abstract>

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a review rebuttal not an abstract

Comment thread data/xml/2026.cmlc.xml
<title>Pop Lyrics through Time: Challenges in Corpus-Based Modeling of Linguistic and Emotional Dynamics in <fixed-case>G</fixed-case>erman Pop Lyrics</title>
<author><first>Roman</first><last>Schneider</last></author>
<pages>32-43</pages>
<abstract>I would like to sincerely thank all reviewers for their time, expertise, and thoughtful feedback on my paper! Please find below my responses to your remarks. ============================================================================ REVIEWER #1 ============================================================================ REMARK 1: Since this paper has been submitted to the CMLC workshop, I would have liked to see more explicit discussions of how the work is linked to challenges in the management of large corpora.... ANSWER: Added a paragraph on this in the conclusion section. REMARK 2: It would also be good to explain how the Songkorpus handle code switching in songs between German, English and other languages... ANSWER: Added a paragraph in Section 2.1. REMARK 3: As new songs are being added, how is IP handled? Does this have to be cleared before lyrics are included in updates to the corpus each year e.g. from the current chart? ANSWER: Added a paragraph in Section 2.1. ============================================================================ REVIEWER #2 ============================================================================ REMARK 1: Section 2.5: The methods used to estimate sentiment intensity and detect MPs are somewhat outdated. However, the human-based evaluation of their quality and reliability is very thorough and convincing. Since manually annotated samples are available, it may be worthwhile to test more modern approaches... ANSWER: Expanded Section 2.4.2 accordingly. REMARK 2: Section 3.2: The conclusions presented in this section are very interesting. It might be useful to examine which pronoun has become more prominent over time... ANSWER: Added a paragraph in Section 3.2. REMARK 3: Examples: I strongly recommend adding concrete examples that illustrate the linguistic features. ANSWER: Added a paragraph in Section 2.3.3. REMARK 4: Figures: The font size in the figures is too small and hard to read. ANSWER: Changed! ============================================================================ REVIEWER #3 ============================================================================ REMARK 1: Could you make more explicit whether the paper’s primary contribution is methodological, substantive, or intentionally balanced between the two? ANSWER: Clarified the contribution in the introduction, then reframed the research questions in Section 1.2. REMARK 2: Could you say a little more about the intended interpretive payoff of features such as pronoun usage and modal particles in this domain? ANSWER: Added a short motivation paragraph in the introduction that directly answers "why these features?", and tried to strengthen interpretive payoff in the Results discussion REMARK 3: How confident can we be in interpreting the results as showing cultural features such as emotional flattening? ANSWER: Please see the new remarks on this in the conclusion.</abstract>

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Another review rebuttal

Comment thread data/xml/2026.cmlc.xml
<author><first>Svetla Peneva</first><last>Koeva</last></author>
<author><first>Ivelina</first><last>Stoyanova</last></author>
<pages>71-75</pages>
<abstract>1. (Reviewer 1) Comment: Some terminology could do with more explanation, such as MARCELL and CURLICAT, however, the general point is very clear. Response: We have added some more details on the international projects that involved the creation of large datasets included in BulNC. 2. (Reviewer 1) Comment: I am especially interested in the enrichment of the corpus with multimodal data. I think this is something that many large corpus managers would be interested in exploring. Response: We added some more details on the multimodal dataset, the organisation of the multimodal data, the ontology description, and the applications. 3. (Reviewer 2) Comment: It does not, however, give the reader a clear idea of current priorities and future directions, and it largely fails to put the work in context of other research. Response: A clarification has been made in the conclusion that we aim at extending the large dataset with more data and extensive metadata description in order to facilitate development of language technologies and fine-tuning of LLMs. 4. (Reviewer 2) Comment: - "Like many other large reference corpora" - please give reference so that readers can know in what context you see your own work Response: We expanded the Introduction with a paragraph citing other related work on large reference corpora: "There are two main approaches to providing search interfaces for large reference corpora. ..." Also, in the conclusion we included more details on large datasets used for LLMs, which provide context for our future work. 5. (Reviewer 2) Comment: - "linguistic and corpus research" - do you see these as two different (sub)disciplines? Explain or use a different wording Response: Thank you for the remark, it is well founded and we reformulated it as ‘linguistic and NLP research’. 6. (Reviewer 2) Comment:- JSONL and CSV - explain what that is and how you are using it for linguistic data – the formats themselves are just very general specifications for textual data. Are you using any particular linguistic standards? Response: More details are provided in the text with respect to the BulNC processing pipelines, and the handling of different file formats. Due to the limited volume of the paper we have not provided details on the linguistic annotation of the corpus. 7. (Reviewer 2) Comment: - "now called the IfGPT dataset" – is that relevant? What does the acronym stand for? Response: In the fourth paragraph of Section 1, we explain: "These efforts led to the development of the large BulNC-based dataset within the project <i>IfGPT: Infrastructure for Fine-tuning Pre-trained Large Language Models</i> (thus, also called the <b>IfGPT dataset</b>), with a special focus on the efficient management of large text data." 8. (Reviewer 2) Comment: - Table 1 - please explain what the different corpus components actually are Response: We clarified in the following way: "Further extensions of the dataset include newly collected and processed texts from various time periods. Older texts, such as news articles, periodicals, and books published before 1990, are also collected and processed using OCR." 9. (Reviewer 2) Comment:- "25 languages" - which ones? How where they chosen? Response: We clarified in the following way: "The selection of languages was based on the availability of wordnets in various languages in the Extended Open Multilingual Wordnet." 10. (Reviewer 2) Comment: - "therefore has a complex graph-based structure" – I don’t see why this follows from the fact that BulNC is designed to support corpus and language research. Can you explain? How does your graph-based structure relate to other approaches to metadata? Response: Justification for the use of Neo4J database is provided: "The metadata are managed using a graph database, Neo4J, that is designed to handle large volumes of interconnected data efficiently and maintains performance under complex queries using the Cypher query language." 11. (Reviewer 2) Comment: - "Corpus Query Tool specifically developed for the BulNC" – reference? URL? Response: Reference is provided both to the BulNC search interface and the IfGPT metadata web search. 12. (Reviewer 2) Comment: - References – these are exclusively self-references. Do you not want to put your work into the context of other CMLC contributions? Response: More references are provided. See 4. 13. (Reviewer 2) Comment: - General remark: You’re leaving implicit what BulNC is *not* doing. Can you devote at least one sentence to your approach to spoken language? What about CMC, learner language, etc.? Response: Currently, we have not extended our work towards including spoken language data or other specialised datasets (e.g., learner data, etc.). 14. (Reviewer 3) Comment: The difference between "BulNC", "BulNC-based dataset" and "IfGPT dataset" is never really made clear and needs to be inferred by the reader. The presentation would benefit greatly from one or two sentences early on that explicitly spell out the relationship between BulNC, the BulNC-based dataset, and the IfGPT dataset. Response: We thank the reviewers for highlighting the need to clarify the relationship between "BulNC", the "BulNC-based dataset", and the "IfGPT dataset". Answer is as 7. above.</abstract>

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

another one. looks like a systematic issue with this workshop

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@mjpost How do you want to handle these? I can correct them in the XML but it may be better to correct at the source in case of a reingestion.

@nschneid

Copy link
Copy Markdown
Collaborator

2026.lrec-1.608: &lt;url&gt;https://hf.co/collections/prachuryyaIITG/aptfiner&lt;/url&gt;

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants