16th Workshop on Building and Using Comparable Corpora

MOTIVATION

In the language engineering and the linguistics communities, research in comparable corpora has been motivated by two main reasons. In language engineering, on the one hand, it is chiefly motivated by the need to use comparable corpora as training data for statistical NLP applications such as statistical and neural machine translation or cross-lingual retrieval. In linguistics, on the other hand, comparable corpora are of interest because they enable cross-language discoveries and comparisons. It is generally accepted in both communities that comparable corpora consist of documents that are comparable in content and form in various degrees and dimensions across several languages. Parallel corpora are on the one end of this spectrum, unrelated corpora on the other.
Comparable corpora have been used in a range of applications, including Information Retrieval, Machine Translation, Cross-lingual text classification, etc. The linguistic definitions and observations related to comparable corpora can improve methods to mine such corpora for applications of statistical NLP, for example to extract parallel corpora from comparable corpora for neural MT. As such, it is of great interest to bring together builders and users of such corpora.

Previous BUCC Workshops

IssueVenue ChairpersonsProceedings
BUCC 2008LREC, Marrakech Pierre Zweigenbaum, Éric Gaussier, Pascale Fung PDF
BUCC 2009ACL, Singapore Pascale Fung, Pierre Zweigenbaum, Reinhard Rapp ACL Anthology page PDF [BibTeX]
BUCC 2010LREC, Valetta Reinhard Rapp, Pierre Zweigenbaum, Serge Sharoff PDF
BUCC 2011ACL, Portland Pierre Zweigenbaum, Reinhard Rapp, Serge Sharoff ACL Anthology page PDF [BibTeX]
BUCC 2012LREC, Istanbul Reinhard Rapp, Marko Tadić, Serge Sharoff, Andrejs Vasiļjevs, Pierre Zweigenbaum PDF
BUCC 2013ACL, Sofia Serge Sharoff, Pierre Zweigenbaum, Reinhard Rapp ACL Anthology page PDF [BibTeX]
BUCC 2014LREC, Reykjavik Pierre Zweigenbaum, Ahmet Aker, Serge Sharoff, Stephan Vogel, Reinhard Rapp PDF [Individual papers]
BUCC 2015ACL, Beijing Pierre Zweigenbaum, Serge Sharoff, Reinhard Rapp ACL Anthology page PDF [BibTeX]
BUCC 2016LREC, Portorož Reinhard Rapp, Pierre Zweigenbaum, Serge Sharoff PDF [Individual papers] [BibTeX]
BUCC 2017ACL, Vancouver Serge Sharoff, Pierre Zweigenbaum, Reinhard Rapp ACL Anthology page PDF [BibTeX]
BUCC 2018LREC, Miyazaki Reinhard Rapp, Pierre Zweigenbaum, Serge Sharoff Proceedings page PDF [Individual papers]
BUCC 2019RANLP, Varna Serge Sharoff, Pierre Zweigenbaum, Reinhard Rapp PDF [Individual papers] [BibTeX]
BUCC 2020LREC, online Reinhard Rapp, Pierre Zweigenbaum, Serge Sharoff ACL Anthology page PDF [BibTeX]
BUCC 2021RANLP, online Reinhard Rapp, Serge Sharoff, Pierre Zweigenbaum ACL Anthology page PDF [BibTeX]
BUCC 2022LREC, Marseille Reinhard Rapp, Pierre Zweigenbaum, Serge Sharoff ACL Anthology page PDF [BibTeX]
BUCC 2023RANLP, Varna Reinhard Rapp, Pierre Zweigenbaum, Serge Sharoff ACL Anthology page PDF [BibTeX]
BUCC 2024LREC-COLING, Torino Pierre Zweigenbaum, Reinhard Rapp, Serge Sharoff ACL Anthology page PDF [BibTeX]
BUCC 2025COLING, Abu Dhabi Serge Sharoff, Ayla Rigouts Terryn, Pierre Zweigenbaum, Reinhard Rapp