18th Workshop on Building and Using Comparable Corpora
MOTIVATION
In the language engineering and linguistics communities, research in comparable corpora has been motivated by two main reasons. In language engineering, on the one hand, it is chiefly motivated by the need to use comparable corpora as training data for statistical NLP applications such as
statistical and neural machine translation or cross-lingual retrieval. In linguistics, on the other hand, comparable corpora are of interest because they enable cross-language discoveries and comparisons. It is generally accepted in both communities that comparable corpora consist of documents
that are comparable in content and form in various degrees and dimensions across several languages. Parallel corpora are on the one end of this spectrum, and unrelated corpora are on the other.
In recent years, the use of comparable corpora for pre-training Large Language Models (LLMs) has led to their impressive multilingual and cross-lingual abilities, which are relevant to a range of applications, including Information Retrieval, Machine Translation, Cross-lingual text
classification, etc. The linguistic definitions and observations related to comparable corpora can improve methods to mine such corpora for applications of statistical NLP, for example, to extract parallel corpora from comparable corpora for neural MT or to improve cross-lingual transfer of
LLMs. As such, it is of great interest to bring together builders and users of such corpora.
Previous BUCC Workshops
Issue | Venue | Chairpersons | Proceedings |
BUCC 2008 | LREC, Marrakech | Pierre Zweigenbaum, Éric Gaussier, Pascale Fung | |
BUCC 2009 | ACL, Singapore | Pascale Fung, Pierre Zweigenbaum, Reinhard Rapp | ACL Anthology page PDF [BibTeX] |
BUCC 2010 | LREC, Valetta | Reinhard Rapp, Pierre Zweigenbaum, Serge Sharoff | |
BUCC 2011 | ACL, Portland | Pierre Zweigenbaum, Reinhard Rapp, Serge Sharoff | ACL Anthology page PDF [BibTeX] |
BUCC 2012 | LREC, Istanbul | Reinhard Rapp, Marko Tadić, Serge Sharoff, Andrejs Vasiļjevs, Pierre Zweigenbaum | |
BUCC 2013 | ACL, Sofia | Serge Sharoff, Pierre Zweigenbaum, Reinhard Rapp | ACL Anthology page PDF [BibTeX] |
BUCC 2014 | LREC, Reykjavik | Pierre Zweigenbaum, Ahmet Aker, Serge Sharoff, Stephan Vogel, Reinhard Rapp | PDF [Individual papers] |
BUCC 2015 | ACL, Beijing | Pierre Zweigenbaum, Serge Sharoff, Reinhard Rapp | ACL Anthology page PDF [BibTeX] |
BUCC 2016 | LREC, Portorož | Reinhard Rapp, Pierre Zweigenbaum, Serge Sharoff | PDF [Individual papers] [BibTeX] |
BUCC 2017 | ACL, Vancouver | Serge Sharoff, Pierre Zweigenbaum, Reinhard Rapp | ACL Anthology page PDF [BibTeX] |
BUCC 2018 | LREC, Miyazaki | Reinhard Rapp, Pierre Zweigenbaum, Serge Sharoff | Proceedings page PDF [Individual papers] |
BUCC 2019 | RANLP, Varna | Serge Sharoff, Pierre Zweigenbaum, Reinhard Rapp | PDF [Individual papers] [BibTeX] |
BUCC 2020 | LREC, online | Reinhard Rapp, Pierre Zweigenbaum, Serge Sharoff | ACL Anthology page PDF [BibTeX] |
BUCC 2021 | RANLP, online | Reinhard Rapp, Serge Sharoff, Pierre Zweigenbaum | ACL Anthology page PDF [BibTeX] |
BUCC 2022 | LREC, Marseille | Reinhard Rapp, Pierre Zweigenbaum, Serge Sharoff | ACL Anthology page PDF [BibTeX] |
BUCC 2023 | RANLP, Varna | Reinhard Rapp, Pierre Zweigenbaum, Serge Sharoff | ACL Anthology page PDF [BibTeX] |
BUCC 2024 | LREC-COLING, Torino | Pierre Zweigenbaum, Reinhard Rapp, Serge Sharoff | ACL Anthology page PDF [BibTeX] |
BUCC 2025 | COLING, Abu Dhabi | Serge Sharoff, Ayla Rigouts Terryn, Pierre Zweigenbaum, Reinhard Rapp |