Ingest LREC and workshops - #9304
Conversation
📊 Ingestion statisticsAnalyzed 47 volume(s):
Top ten authors by paper counts: Wajdi Zaghouani (21), Georg Rehm (11), Els Lefever (9), James Pustejovsky (9), Jelke Bloem (9), Barbara Plank (8), Eiji Aramaki (7), Hafsteinn Einarsson (7), Marco Carlo Passarotti (7), Maxime Amblard (7) Papers by number of authorsxychart-beta
title "Papers by number of authors"
x-axis "Number of authors" [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 21, 26, 29, 30, 78, 107]
y-axis "Papers" 0 --> 454
bar [156, 369, 454, 348, 213, 133, 80, 42, 29, 24, 9, 2, 9, 4, 4, 3, 1, 3, 1, 1, 2, 1, 1, 1]
|
|
Would it be easy to rerun with a temporary fix for the newline bug? I see ".The" in the data, for example. |
ingest.py is now updated on master. |
Okay, I re-ingested LREC main, and there were two small changes. It was a bunch of work correcting mistakes, so I'd prefer to just apply this going forward from now on (and not do this for all workshops). |
|
@nschneid Something else that has come up, is they have asked for us to point to the LREC-hosted PDFs. We still have our local copies. Ideally we could point to both but I think at the moment we have to choose one. |
|
It seems that starting with LREC 2016 we have been hosting PDFs. The advantage of that is more redundancy (in case one of the servers is down). The advantage of pointing to theirs would be that we would not have to worry about PDF revisions on our end. Though of course we would want to know about title/abstract/author changes or retractions. |
Hmm, then maybe |
It is. Input file cdrom/bib/2026.lrec2026-1.938.bib:
|
|
OK. Based on some stats I have a hunch that the newline bug was mainly affecting aclpub2 proceedings. Maybe the LREC undersegmentations are just user typos. |
|
When your work on the XML is done I can run my script from #9345 as that is currently independent of the ingestion code. |
| <author><first>Svetla Peneva</first><last>Koeva</last></author> | ||
| <author><first>Ivelina</first><last>Stoyanova</last></author> | ||
| <pages>12-24</pages> | ||
| <abstract>Many thanks to all reviewers for the detailed comments and suggestions. Misspelling, formatting errors and other minor issues have all been corrected, and are not listed below. 1. (Reviewer 1) Comment: For example, the relation of some described corpora to the Bulgarian National Corpus (like MIC21, Bulgarian MARCELL, General News in Bulgarian, ...) is not clear to me. Are they part of BulNC or the other way around? It’s stated that the BulNC is part of CURLICAT, so I would assume the relation with the other corpora is similar? Response: The conclusion now clarifies the relations between BulNC and IfGPT and its subsets. The paper was also restructured to clarify these issues. 2. (Reviewer 1) Comment: Some sections on the paper reference related work, where again the relation is not clear to me, like on page 3 "Another direction, which still presents significant challenges, is towards multilingual data. O’Keeffe et al. (2024) describe a pipeline for capturing professional video-call interactions, including screen recordings, speaker tracking, and facial expression data. Macaire et al. (2024) present a speech text pictogram corpus for French (230 hours), targeted at augmentative and alternative communication research. Lai and Pustejovsky (2024) develop an annotation scheme for iconic, deictic, and beat gestures anchored in Abstract Meaning Representation." As these publications follow the MIC21 corpus, it’s not clear to me, how they relate. Response: These references are reduced; a citation for the MIC corpus is provided where more details are available on the related works of MIC. 3. (Reviewer 1) Comment: I would be very interested in the benefits of a Graph database for the Metadata and I think it would be very valuable to describe, why the Graph database is more appropriate for Metadata than commonly used alternatives. The described relations are not totally convincing to me. Response: A paragraph is added in the Metadata management section on the justification of the use of a graph database. 4. (Reviewer 1) Comment: The publicly accessible web interface should be linked to. It’s also not clear to me, if it provides a fulltext search or only a keyword search in the Metadata (i.e. in the Graph database). Response: Links are provided to both the search interface of BulNC (full text search, mainly for linguistic research) and to IfGPT metadata search interface (allowing selection of subdatasets for NLP tasks and LLM fine-tuning). 5. (Reviewer 2) Comment: My first doubt is why all the data is presented as part or at least related to the Bulgarian National Corpus - to me, national corpora are reference language corpora with all the characteristics this entails, like a carefully balanced corpus representative of contemporary standard Bulgarian. I would not expect to see in this context mentioned artificially produced language for LLMs, multilingual or foreign language corpora or image corpora. To me it would make a lot more sense the re-cast the paper as presenting (newly) available language resources for Bulgarian (maybe in the context of the Bulgarian CLARIN / CLARIAH, which is not even mentioned!) rather than shoe-horning them to the BulNC. Response: The ties between the BulNC and the large dataset IfGPT has been clarified in the Introduction, in the text and in the Conclusion. 6. (Reviewer 2) Comment: Second, it would help the reader that, rather than just listing all the resources, a table with their key characteristics would be provided first, and then each introduced. Similarly, the endless repetitions of "so-and-so many JSON files" for each domain in Sec. 2.5 are not really helpful, as, first, texts or words are a better metric than files, and, second, this information would also be better presented in a table. Response: A table has now been provided listing the resources and their key properties, including size. 7. (Reviewer 2) Comment: "The BulNC-based dataset is publicly accessible through a dedicated web interface" I don’t see that these datasets are in any way BulNC-based. Response: Clarifications are made on this issue in the Introduction and the Conclusion. 8. (Reviewer 3) Comment: I am missing (at least a sketchy) description of tools used for tagging and/or parsing the corpus data. Response: A reference to the Bulgarian Language Processing Tool Set is now provided. 9. (Reviewer 3) Comment: It is also not clear why the corpus is still using the web interface developed before 2014, and not any newer tool (with a FLOSS license, such as CQPweb or NoSketch Engine). Response: This has been clarified in section 2.1. 10. (Reviewer 3) Comment: I would also expect the text to include summary statistics (preferably in a single table), so the sizes of the respective resources can be compared. Response: Table 1 is now provided for this purpose.</abstract> |
There was a problem hiding this comment.
This is a review rebuttal not an abstract
| <title>Pop Lyrics through Time: Challenges in Corpus-Based Modeling of Linguistic and Emotional Dynamics in <fixed-case>G</fixed-case>erman Pop Lyrics</title> | ||
| <author><first>Roman</first><last>Schneider</last></author> | ||
| <pages>32-43</pages> | ||
| <abstract>I would like to sincerely thank all reviewers for their time, expertise, and thoughtful feedback on my paper! Please find below my responses to your remarks. ============================================================================ REVIEWER #1 ============================================================================ REMARK 1: Since this paper has been submitted to the CMLC workshop, I would have liked to see more explicit discussions of how the work is linked to challenges in the management of large corpora.... ANSWER: Added a paragraph on this in the conclusion section. REMARK 2: It would also be good to explain how the Songkorpus handle code switching in songs between German, English and other languages... ANSWER: Added a paragraph in Section 2.1. REMARK 3: As new songs are being added, how is IP handled? Does this have to be cleared before lyrics are included in updates to the corpus each year e.g. from the current chart? ANSWER: Added a paragraph in Section 2.1. ============================================================================ REVIEWER #2 ============================================================================ REMARK 1: Section 2.5: The methods used to estimate sentiment intensity and detect MPs are somewhat outdated. However, the human-based evaluation of their quality and reliability is very thorough and convincing. Since manually annotated samples are available, it may be worthwhile to test more modern approaches... ANSWER: Expanded Section 2.4.2 accordingly. REMARK 2: Section 3.2: The conclusions presented in this section are very interesting. It might be useful to examine which pronoun has become more prominent over time... ANSWER: Added a paragraph in Section 3.2. REMARK 3: Examples: I strongly recommend adding concrete examples that illustrate the linguistic features. ANSWER: Added a paragraph in Section 2.3.3. REMARK 4: Figures: The font size in the figures is too small and hard to read. ANSWER: Changed! ============================================================================ REVIEWER #3 ============================================================================ REMARK 1: Could you make more explicit whether the paper’s primary contribution is methodological, substantive, or intentionally balanced between the two? ANSWER: Clarified the contribution in the introduction, then reframed the research questions in Section 1.2. REMARK 2: Could you say a little more about the intended interpretive payoff of features such as pronoun usage and modal particles in this domain? ANSWER: Added a short motivation paragraph in the introduction that directly answers "why these features?", and tried to strengthen interpretive payoff in the Results discussion REMARK 3: How confident can we be in interpreting the results as showing cultural features such as emotional flattening? ANSWER: Please see the new remarks on this in the conclusion.</abstract> |
| <author><first>Svetla Peneva</first><last>Koeva</last></author> | ||
| <author><first>Ivelina</first><last>Stoyanova</last></author> | ||
| <pages>71-75</pages> | ||
| <abstract>1. (Reviewer 1) Comment: Some terminology could do with more explanation, such as MARCELL and CURLICAT, however, the general point is very clear. Response: We have added some more details on the international projects that involved the creation of large datasets included in BulNC. 2. (Reviewer 1) Comment: I am especially interested in the enrichment of the corpus with multimodal data. I think this is something that many large corpus managers would be interested in exploring. Response: We added some more details on the multimodal dataset, the organisation of the multimodal data, the ontology description, and the applications. 3. (Reviewer 2) Comment: It does not, however, give the reader a clear idea of current priorities and future directions, and it largely fails to put the work in context of other research. Response: A clarification has been made in the conclusion that we aim at extending the large dataset with more data and extensive metadata description in order to facilitate development of language technologies and fine-tuning of LLMs. 4. (Reviewer 2) Comment: - "Like many other large reference corpora" - please give reference so that readers can know in what context you see your own work Response: We expanded the Introduction with a paragraph citing other related work on large reference corpora: "There are two main approaches to providing search interfaces for large reference corpora. ..." Also, in the conclusion we included more details on large datasets used for LLMs, which provide context for our future work. 5. (Reviewer 2) Comment: - "linguistic and corpus research" - do you see these as two different (sub)disciplines? Explain or use a different wording Response: Thank you for the remark, it is well founded and we reformulated it as ‘linguistic and NLP research’. 6. (Reviewer 2) Comment:- JSONL and CSV - explain what that is and how you are using it for linguistic data – the formats themselves are just very general specifications for textual data. Are you using any particular linguistic standards? Response: More details are provided in the text with respect to the BulNC processing pipelines, and the handling of different file formats. Due to the limited volume of the paper we have not provided details on the linguistic annotation of the corpus. 7. (Reviewer 2) Comment: - "now called the IfGPT dataset" – is that relevant? What does the acronym stand for? Response: In the fourth paragraph of Section 1, we explain: "These efforts led to the development of the large BulNC-based dataset within the project <i>IfGPT: Infrastructure for Fine-tuning Pre-trained Large Language Models</i> (thus, also called the <b>IfGPT dataset</b>), with a special focus on the efficient management of large text data." 8. (Reviewer 2) Comment: - Table 1 - please explain what the different corpus components actually are Response: We clarified in the following way: "Further extensions of the dataset include newly collected and processed texts from various time periods. Older texts, such as news articles, periodicals, and books published before 1990, are also collected and processed using OCR." 9. (Reviewer 2) Comment:- "25 languages" - which ones? How where they chosen? Response: We clarified in the following way: "The selection of languages was based on the availability of wordnets in various languages in the Extended Open Multilingual Wordnet." 10. (Reviewer 2) Comment: - "therefore has a complex graph-based structure" – I don’t see why this follows from the fact that BulNC is designed to support corpus and language research. Can you explain? How does your graph-based structure relate to other approaches to metadata? Response: Justification for the use of Neo4J database is provided: "The metadata are managed using a graph database, Neo4J, that is designed to handle large volumes of interconnected data efficiently and maintains performance under complex queries using the Cypher query language." 11. (Reviewer 2) Comment: - "Corpus Query Tool specifically developed for the BulNC" – reference? URL? Response: Reference is provided both to the BulNC search interface and the IfGPT metadata web search. 12. (Reviewer 2) Comment: - References – these are exclusively self-references. Do you not want to put your work into the context of other CMLC contributions? Response: More references are provided. See 4. 13. (Reviewer 2) Comment: - General remark: You’re leaving implicit what BulNC is *not* doing. Can you devote at least one sentence to your approach to spoken language? What about CMC, learner language, etc.? Response: Currently, we have not extended our work towards including spoken language data or other specialised datasets (e.g., learner data, etc.). 14. (Reviewer 3) Comment: The difference between "BulNC", "BulNC-based dataset" and "IfGPT dataset" is never really made clear and needs to be inferred by the reader. The presentation would benefit greatly from one or two sentences early on that explicitly spell out the relationship between BulNC, the BulNC-based dataset, and the IfGPT dataset. Response: We thank the reviewers for highlighting the need to clarify the relationship between "BulNC", the "BulNC-based dataset", and the "IfGPT dataset". Answer is as 7. above.</abstract> |
There was a problem hiding this comment.
another one. looks like a systematic issue with this workshop
There was a problem hiding this comment.
@mjpost How do you want to handle these? I can correct them in the XML but it may be better to correct at the source in case of a reingestion.
|
2026.lrec-1.608: |
Administrative checklist:
masterbranch<meta>block<venue>ws</venue>tag to its meta block<event>blockNavigate to the preview site and check the following:
After the PR is closed, for all events:
YYYY-MM-DD-{event}Create followup issues for the following tasks (usually just for ACL events)