Publications

2026

 
Clustering Analysis for Error Detection in Named Entity Recognition Datasets. Matthew Flynn, Timothy Obiso, Sam Newman, and Constantine Lignos. In Yang Janet Liu and Luke Gessler, editors, Proceedings of the 20th Linguistic Annotation Workshop (LAW XX), pages 229–240, San Diego, California, USA, July 2026. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

This paper introduces a method for the automatic detection of annotation errors and corrections in named entity recognition datasets using a novel two-stage dimension reduction of dense sentence embeddings. We first find the top-n principal components of an embedding and then use UMAP for second-stage, non-linear dimension reduction and clustering using different distance metrics. We analyze these clusters using silhouette scores to flag outlier mentions for correction. Using the corrections in the CoNLL# dataset as a benchmark, all of the top-five outliers needed correction, as did 7 of the top-10. This approach also identified 32 of the top-50 outlier mentions that are corrections. This method offers a relatively low-effort way to leverage text embeddings and dimensionality reduction to identify likely annotation errors. We release related code and data at https://github.com/bltlab/clustering-for-ner.

 
How multilingual are multilingual LLMs? A case study in Northern Sámi-Finnish Translation. Jonne Sälevä and Constantine Lignos. In Hansi Hettiarachchi, Tharindu Ranasinghe, Alistair Plum, Paul Rayson, Ruslan Mitkov, Mohamed Gaber, Damith Premasiri, Fiona Anting Tan, and Lasitha Uyangodage, editors, Proceedings of the Second Workshop on Language Models for Low-Resource Languages (LoResLM 2026), pages 484–492, Rabat, Morocco, March 2026. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

We use Finnish and Northern Sámi as a case study to investigate how suitable multilingual LLMs are for low-resource machine translation and how much performance can be improved using supervised finetuning with varying amounts of parallel data. Our experiments on zero-shot translation reveal that mainstream multilingual LLMs from a variety of model families are unsuitable for translation between our chosen languages as-is, regardless of the generation hyperparameters. On the other hand, our experiments on supervised finetuning reveal that even relatively small amounts of parallel data can be very useful for improving performance in both translation directions.

 
Proceedings of the 7th Workshop on African Natural Language Processing (AfricaNLP 2026). Everlyn Asiko Chimoto, Constantine Lignos, Shamsuddeen Muhammad, Idris Abdulmumin, Clemencia Siro, and David Ifeoluwa Adelani, editors, Rabat, Morocco, March 2026. Association for Computational Linguistics. [ bib | DOI | http ]

2025

 
OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages. Chester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä, and Constantine Lignos. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng, editors, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 33649–33674, Suzhou, China, November 2025. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

We present OpenNER 1.0, a standardized collection of openly-available named entity recognition (NER) datasets.OpenNER contains 36 NER corpora that span 52 languages, human-annotated in varying named entity ontologies.We correct annotation format issues, standardize the original datasets into a uniform representation with consistent entity type names across corpora, and provide the collection in a structure that enables research in multilingual and multi-ontology NER.We provide baseline results using three pretrained multilingual language models and two large language models to compare the performance of recent models and facilitate future research in NER.We find that no single model is best in all languages and that significant work remains to obtain high performance from LLMs on the NER task.OpenNER is released at https://github.com/bltlab/open-ner.

 
Beyond statistical significance: Quantifying uncertainty and statistical variability in multilingual and multitask NLP evaluation. Jonne Sälevä, Duygu Ataman, and Constantine Lignos. In Kentaro Inui, Sakriani Sakti, Haofen Wang, Derek F. Wong, Pushpak Bhattacharyya, Biplab Banerjee, Asif Ekbal, Tanmoy Chakraborty, and Dhirendra Pratap Singh, editors, Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics, pages 2304–2321, Mumbai, India, December 2025. The Asian Federation of Natural Language Processing and The Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

We introduce a set of resampling-based methods for quantifying uncertainty and statistical precision of evaluation metrics in multilingual and/or multitask NLP benchmarks.We show how experimental variation in performance scores arises from both model and data-related sources, and that accounting for both of them is necessary to avoid substantially underestimating the overall variability over hypothetical replications.Using multilingual question answering, machine translation, and named entity recognition as example tasks, we also demonstrate how resampling methods are useful for quantifying the replication uncertainty of various quantities used in leaderboards such as model rankings and pairwise differences between models.

 
Proceedings of the Sixth Workshop on African Natural Language Processing (AfricaNLP 2025). Constantine Lignos, Idris Abdulmumin, and David Adelani, editors, Vienna, Austria, July 2025. Association for Computational Linguistics. [ bib | DOI | http ]
 
MetaMeme: A Dataset for Meme Template and Meta-Category Classification. Benjamin Lambright, Jordan Youner, and Constantine Lignos. In Abteen Ebrahimi, Samar Haider, Emmy Liu, Sammar Haider, Maria Leonor Pacheco, and Shira Wein, editors, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 4: Student Research Workshop), pages 356–367, Albuquerque, USA, April 2025. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

This paper introduces a new dataset for classifying memes by their template and communicative intent.It includes a broad selection of meme templates and examples scraped from imgflip and a smaller hand-annotated set of memes scraped from Reddit.The Reddit memes have been annotated for meta-category using a novel annotation scheme that classifies memes by the structure of the perspective they are being used to communicate.YOLOv11 and ChatGPT 4o are used to provide baseline modeling results.We find that YOLO struggles with template classification on real-world data but outperforms ChatGPT in classifying meta-categories.

2024

 
Language Model Priors and Data Augmentation Strategies for Low-resource Machine Translation: A Case Study Using Finnish to Northern Sámi. Jonne Sälevä and Constantine Lignos. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Association for Computational Linguistics: ACL 2024, pages 12949–12956, Bangkok, Thailand, August 2024. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

We investigate ways of using monolingual data in both the source and target languages for improving low-resource machine translation. As a case study, we experiment with translation from Finnish to Northern Sámi.Our experiments show that while conventional backtranslation remains a strong contender, using synthetic target-side data when training backtranslation models can be helpful as well.We also show that monolingual data can be used to train a language model which can act as a regularizer without any augmentation of parallel data.

 
CoNLL#: Fine-grained Error Analysis and a Corrected Test Set for CoNLL-03 English. Andrew Rueda, Elena Alvarez-Mellado, and Constantine Lignos. In Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue, editors, Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 3718–3728, Torino, Italia, May 2024. ELRA and ICCL. [ bib | http ]
Abstract

Modern named entity recognition systems have steadily improved performance in the age of larger and more powerful neural models. However, over the past several years, the state-of-the-art has seemingly hit another plateau on the benchmark CoNLL-03 English dataset. In this paper, we perform a deep dive into the test outputs of the highest-performing NER models, conducting a fine-grained evaluation of their performance by introducing new document-level annotations on the test set. We go beyond F1 scores by categorizing errors in order to interpret the true state of the art for NER and guide future work. We review previous attempts at correcting the various flaws of the test set and introduce CoNLL#, a new corrected version of the test set that addresses its systematic and most prevalent errors, allowing for low-noise, interpretable error analysis.

 
ParaNames 1.0: Creating an Entity Name Corpus for 400+ Languages Using Wikidata. Jonne Sälevä and Constantine Lignos. In Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue, editors, Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 12599–12610, Torino, Italia, May 2024. ELRA and ICCL. [ bib | http ]
Abstract

We introduce ParaNames, a massively multilingual parallel name resource consisting of 140 million names spanning over 400 languages. Names are provided for 16.8 million entities, and each entity is mapped from a complex type hierarchy to a standard type (PER/LOC/ORG). Using Wikidata as a source, we create the largest resource of this type to date. We describe our approach to filtering and standardizing the data to provide the best quality possible. ParaNames is useful for multilingual language processing, both in defining tasks for name translation/transliteration and as supplementary data for tasks such as named entity recognition and linking. We demonstrate the usefulness of ParaNames on two tasks. First, we perform canonical name translation between English and 17 other languages. Second, we use it as a gazetteer for multilingual named entity recognition, obtaining performance improvements on all 10 languages evaluated.

 
QueryNER: Segmentation of E-commerce Queries. Chester Palen-Michel, Lizzie Liang, Zhe Wu, and Constantine Lignos. In Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue, editors, Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 13455–13470, Torino, Italia, May 2024. ELRA and ICCL. [ bib | http ]
Abstract

We present QueryNER, a manually-annotated dataset and accompanying model for e-commerce query segmentation. Prior work in sequence labeling for e-commerce has largely addressed aspect-value extraction which focuses on extracting portions of a product title or query for narrowly defined aspects. Our work instead focuses on the goal of dividing a query into meaningful chunks with broadly applicable types. We report baseline tagging results and conduct experiments comparing token and entity dropping for null and low recall query recovery. Challenging test sets are created using automatic transformations and show how simple data augmentation techniques can make the models more robust to noise. We make the QueryNER dataset publicly available.

2023

 
Findings of the CoCo4MT 2023 Shared Task on Corpus Construction for Machine Translation. Ananya Ganesh, Marine Carpuat, William Chen, Katharina Kann, Constantine Lignos, John E. Ortega, Jonne Saleva, Shabnam Tafreshi, and Rodolfo Zevallos. In Proceedings of the Second Workshop on Corpus Generation and Corpus Augmentation for Machine Translation, pages 22–27, Macau SAR, China, September 2023. Asia-Pacific Association for Machine Translation. [ bib | http ]
 
LR-Sum: Summarization for Less-Resourced Languages. Chester Palen-Michel and Constantine Lignos. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Findings of the Association for Computational Linguistics: ACL 2023, pages 6829–6844, Toronto, Canada, July 2023. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

We introduce LR-Sum, a new permissively-licensed dataset created with the goal of enabling further research in automatic summarization for less-resourced languages.LR-Sum contains human-written summaries for 40 languages, many of which are less-resourced. We describe our process for extracting and filtering the dataset from the Multilingual Open Text corpus (Palen-Michel et al., 2022).The source data is public domain newswire collected from from Voice of America websites, and LR-Sum is released under a Creative Commons license (CC BY 4.0), making it one of the most openly-licensed multilingual summarization datasets. We describe abstractive and extractive summarization experiments to establish baselines and discuss the limitations of this dataset.

 
What changes when you randomly choose BPE merge operations? Not much. Jonne Saleva and Constantine Lignos. In Shabnam Tafreshi, Arjun Akula, João Sedoc, Aleksandr Drozd, Anna Rogers, and Anna Rumshisky, editors, Proceedings of the Fourth Workshop on Insights from Negative Results in NLP, pages 59–66, Dubrovnik, Croatia, May 2023. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

We introduce two simple randomized variants of byte pair encoding (BPE) and explore whether randomizing the selection of merge operations substantially affects a downstream machine translation task. We focus on translation into morphologically rich languages, hypothesizing that this task may show sensitivity to the method of choosing subwords. Analysis using a Bayesian linear model indicates that one variant performs nearly indistinguishably compared to standard BPE while the other degrades performance less than we anticipated. We conclude that although standard BPE is widely used, there exists an interesting universe of potential variations on it worth investigating. Our code is available at: https://github.com/bltlab/random-bpe.

 
Improving NER Research Workflows with SeqScore. Constantine Lignos, Maya Kruse, and Andrew Rueda. In Liling Tan, Dmitrijs Milajevs, Geeticka Chauhan, Jeremy Gwinnup, and Elijah Rippeth, editors, Proceedings of the 3rd Workshop for Natural Language Processing Open Source Software (NLP-OSS 2023), pages 147–152, Singapore, December 2023. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

We describe the features of SeqScore, an MIT-licensed Python toolkit for working with named entity recognition (NER) data.While SeqScore began as a tool for NER scoring, it has been expanded to help with the full lifecycle of working with NER data: validating annotation, providing at-a-glance and detailed summaries of the data, modifying annotation to support experiments, scoring system output, and aiding with error analysis.SeqScore is released via PyPI (https://pypi.org/project/seqscore/) and development occurs on GitHub (https://github.com/bltlab/seqscore).

2022

 
ParaNames: A Massively Multilingual Entity Name Corpus. Jonne Sälevä and Constantine Lignos. In Ekaterina Vylomova, Edoardo Ponti, and Ryan Cotterell, editors, Proceedings of the 4th Workshop on Research in Computational Linguistic Typology and Multilingual NLP, pages 103–105, Seattle, Washington, July 2022. Association for Computational Linguistics. [ bib | DOI | http ]
 
Detecting Unassimilated Borrowings in Spanish: An Annotated Corpus and Approaches to Modeling. Elena Álvarez-Mellado and Constantine Lignos. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3868–3888, Dublin, Ireland, May 2022. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

This work presents a new resource for borrowing identification and analyzes the performance and errors of several models on this task. We introduce a new annotated corpus of Spanish newswire rich in unassimilated lexical borrowings---words from one language that are introduced into another without orthographic adaptation---and use it to evaluate how several sequence labeling models (CRF, BiLSTM-CRF, and Transformer-based models) perform. The corpus contains 370,000 tokens and is larger, more borrowing-dense, OOV-rich, and topic-varied than previous corpora available for this task. Our results show that a BiLSTM-CRF model fed with subword embeddings along with either Transformer-based embeddings pretrained on codeswitched data or a combination of contextualized word embeddings outperforms results obtained by a multilingual BERT-based model.

 
MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition. David Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani, Michael Beukman, Chester Palen-Michel, Constantine Lignos, Jesujoba Alabi, Shamsuddeen Muhammad, Peter Nabende, Cheikh M. Bamba Dione, Andiswa Bukula, Rooweither Mabuya, Bonaventure F. P. Dossou, Blessing Sibanda, Happy Buzaaba, Jonathan Mukiibi, Godson Kalipe, Derguene Mbaye, Amelia Taylor, Fatoumata Kabore, Chris Chinenye Emezue, Anuoluwapo Aremu, Perez Ogayo, Catherine Gitau, Edwin Munkoh-Buabeng, Victoire Memdjokam Koagne, Allahsera Auguste Tapo, Tebogo Macucwa, Vukosi Marivate, Mboning Tchiaze Elvis, Tajuddeen Gwadabe, Tosin Adewumi, Orevaoghene Ahia, and Joyce Nakatumba-Nabende. In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 4488–4508, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

African languages are spoken by over a billion people, but they are under-represented in NLP research and development. Multiple challenges exist, including the limited availability of annotated training and evaluation datasets as well as the lack of understanding of which settings, languages, and recently proposed methods like cross-lingual transfer will be effective. In this paper, we aim to move towards solutions for these challenges, focusing on the task of named entity recognition (NER). We present the creation of the largest to-date human-annotated NER dataset for 20 African languages. We study the behaviour of state-of-the-art cross-lingual transfer methods in an Africa-centric setting, empirically demonstrating that the choice of source transfer language significantly affects performance. While much previous work defaults to using English as the source language, our results show that choosing the best transfer language improves zero-shot F1 scores by an average of 14% over 20 languages as compared to using English.

 
Proceedings of the 15th biennial conference of the Association for Machine Translation in the Americas (Workshop 2: Corpus Generation and Corpus Augmentation for Machine Translation). John E. Ortega, Marine Carpuat, William Chen, Katharina Kann, Constantine Lignos, Maja Popovic, and Shabnam Tafreshi, editors. Association for Machine Translation in the Americas, September 2022. [ bib | http ]
 
Toward More Meaningful Resources for Lower-resourced Languages. Constantine Lignos, Nolan Holley, Chester Palen-Michel, and Jonne Sälevä. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Findings of the Association for Computational Linguistics: ACL 2022, pages 523–532, Dublin, Ireland, May 2022. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

In this position paper, we describe our perspective on how meaningful resources for lower-resourced languages should be developed in connection with the speakers of those languages. Before advancing that position, we first examine two massively multilingual resources used in language technology development, identifying shortcomings that limit their usefulness. We explore the contents of the names stored in Wikidata for a few lower-resourced languages and find that many of them are not in fact in the languages they claim to be, requiring non-trivial effort to correct. We discuss quality issues present in WikiAnn and evaluate whether it is a useful supplement to hand-annotated data. We then discuss the importance of creating annotations for lower-resourced languages in a thoughtful and ethical way that includes the language speakers as part of the development process. We conclude with recommended guidelines for resource development.

 
Multilingual Open Text Release 1: Public Domain News in 44 Languages. Chester Palen-Michel, June Kim, and Constantine Lignos. In Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Jan Odijk, and Stelios Piperidis, editors, Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 2080–2089, Marseille, France, June 2022. European Language Resources Association. [ bib | http ]
Abstract

We present a Multilingual Open Text (MOT), a new multilingual corpus containing text in 44 languages, many of which have limited existing text resources for natural language processing. The first release of the corpus contains over 2.8 million news articles and an additional 1 million short snippets (photo captions, video descriptions, etc.) published between 2001–2022 and collected from Voice of America`s news websites. We describe our process for collecting, filtering, and processing the data. The source material is in the public domain, our collection is licensed using a creative commons license (CC BY 4.0), and all software used to create the corpus is released under the MIT License. The corpus will be regularly updated as additional documents are published.

 
Borrowing or Codeswitching? Annotating for Finer-Grained Distinctions in Language Mixing. Elena Alvarez-Mellado and Constantine Lignos. In Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Jan Odijk, and Stelios Piperidis, editors, Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 3195–3201, Marseille, France, June 2022. European Language Resources Association. [ bib | http ]
Abstract

We present a new corpus of Twitter data annotated for codeswitching and borrowing between Spanish and English. The corpus contains 9,500 tweets annotated at the token level with codeswitches, borrowings, and named entities. This corpus differs from prior corpora of codeswitching in that we attempt to clearly define and annotate the boundary between codeswitching and borrowing and do not treat common internet-speak (lol, etc.) as codeswitching when used in an otherwise monolingual context. The result is a corpus that enables the study and modeling of Spanish-English borrowing and codeswitching on Twitter in one dataset. We present baseline scores for modeling the labels of this corpus using Transformer-based language models. The annotation itself is released with a CC BY 4.0 license, while the text it applies to is distributed in compliance with the Twitter terms of service.

 
Proceedings of the Workshop on Dataset Creation for Lower-Resourced Languages within the 13th Language Resources and Evaluation Conference. Jonne Sälevä and Constantine Lignos, editors, Marseille, France, June 2022. European Language Resources Association. [ bib | http ]

2021

 
TMR: Evaluating NER Recall on Tough Mentions. Jingxuan Tu and Constantine Lignos. In Ionut-Teodor Sorodoc, Madhumita Sushil, Ece Takmaz, and Eneko Agirre, editors, Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop, pages 155–163, Online, April 2021. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

We propose the Tough Mentions Recall (TMR) metrics to supplement traditional named entity recognition (NER) evaluation by examining recall on specific subsets of tough mentions: unseen mentions, those whose tokens or token/type combination were not observed in training, and type-confusable mentions, token sequences with multiple entity types in the test data. We demonstrate the usefulness of these metrics by evaluating corpora of English, Spanish, and Dutch using five recent neural architectures. We identify subtle differences between the performance of BERT and Flair on two English NER corpora and identify a weak spot in the performance of current models in Spanish. We conclude that the TMR metrics enable differentiation between otherwise similar-scoring systems and identification of patterns in performance that would go unnoticed from overall precision, recall, and F1.

 
The Effectiveness of Morphology-aware Segmentation in Low-Resource Neural Machine Translation. Jonne Saleva and Constantine Lignos. In Ionut-Teodor Sorodoc, Madhumita Sushil, Ece Takmaz, and Eneko Agirre, editors, Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop, pages 164–174, Online, April 2021. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

This paper evaluates the performance of several modern subword segmentation methods in a low-resource neural machine translation setting. We compare segmentations produced by applying BPE at the token or sentence level with morphologically-based segmentations from LMVR and MORSEL. We evaluate translation tasks between English and each of Nepali, Sinhala, and Kazakh, and predict that using morphologically-based segmentation methods would lead to better performance in this setting. However, comparing to BPE, we find that no consistent and reliable differences emerge between the segmentation methods. While morphologically-based methods outperform BPE in a few cases, what performs best tends to vary across tasks, and the performance of segmentation methods is often statistically indistinguishable.

 
SeqScore: Addressing Barriers to Reproducible Named Entity Recognition Evaluation. Chester Palen-Michel, Nolan Holley, and Constantine Lignos. In Yang Gao, Steffen Eger, Wei Zhao, Piyawat Lertvittayakumjorn, and Marina Fomicheva, editors, Proceedings of the 2nd Workshop on Evaluation and Comparison of NLP Systems, pages 40–50, Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

To address a looming crisis of unreproducible evaluation for named entity recognition, we propose guidelines and introduce SeqScore, a software package to improve reproducibility. The guidelines we propose are extremely simple and center around transparency regarding how chunks are encoded and scored. We demonstrate that despite the apparent simplicity of NER evaluation, unreported differences in the scoring procedure can result in changes to scores that are both of noticeable magnitude and statistically significant. We describe SeqScore, which addresses many of the issues that cause replication failures.

 
Macro-Average: Rare Types Are Important Too. Thamme Gowda, Weiqiu You, Constantine Lignos, and Jonathan May. In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tur, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou, editors, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1138–1157, Online, June 2021. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

While traditional corpus-level evaluation metrics for machine translation (MT) correlate well with fluency, they struggle to reflect adequacy. Model-based MT metrics trained on segment-level human judgments have emerged as an attractive replacement due to strong correlation results. These models, however, require potentially expensive re-training for new domains and languages. Furthermore, their decisions are inherently non-transparent and appear to reflect unwelcome biases. We explore the simple type-based classifier metric, MacroF1, and study its applicability to MT evaluation. We find that MacroF1 is competitive on direct assessment, and outperforms others in indicating downstream cross-lingual information retrieval task performance. Further, we show that MacroF1 can be used to effectively compare supervised and unsupervised neural machine translation, and reveal significant qualitative differences in the methods' outputs.

 
MasakhaNER: Named Entity Recognition for African Languages. David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D'souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, Stephen Mayhew, Israel Abebe Azime, Shamsuddeen H. Muhammad, Chris Chinenye Emezue, Joyce Nakatumba-Nabende, Perez Ogayo, Aremu Anuoluwapo, Catherine Gitau, Derguene Mbaye, Jesujoba Alabi, Seid Muhie Yimam, Tajuddeen Rabiu Gwadabe, Ignatius Ezeani, Rubungo Andre Niyongabo, Jonathan Mukiibi, Verrah Otiende, Iroro Orife, Davis David, Samba Ngom, Tosin Adewumi, Paul Rayson, Mofetoluwa Adeyemi, Gerald Muriuki, Emmanuel Anebi, Chiamaka Chukwuneke, Nkiruka Odu, Eric Peter Wairagala, Samuel Oyerinde, Clemencia Siro, Tobius Saul Bateesa, Temilola Oloyede, Yvonne Wambui, Victor Akinode, Deborah Nabagereka, Maurice Katusiime, Ayodele Awokoya, Mouhamadane MBOUP, Dibora Gebreyohannes, Henok Tilaye, Kelechi Nwaike, Degaga Wolde, Abdoulaye Faye, Blessing Sibanda, Orevaoghene Ahia, Bonaventure F. P. Dossou, Kelechi Ogueji, Thierno Ibrahima DIOP, Abdoulaye Diallo, Adewale Akinfaderin, Tendai Marengereke, and Salomey Osei. Transactions of the Association for Computational Linguistics, 9:1116–1131, 2021. [ bib | DOI | http ]
Abstract

We take a step towards addressing the under- representation of the African continent in NLP research by bringing together different stakeholders to create the first large, publicly available, high-quality dataset for named entity recognition (NER) in ten African languages. We detail the characteristics of these languages to help researchers and practitioners better understand the challenges they pose for NER tasks. We analyze our datasets and conduct an extensive empirical evaluation of state- of-the-art methods across both supervised and transfer learning settings. Finally, we release the data, code, and models to inspire future research on African NLP.1

2020

 
If You Build Your Own NER Scorer, Non-replicable Results Will Come. Constantine Lignos and Marjan Kamyab. In Anna Rogers, João Sedoc, and Anna Rumshisky, editors, Proceedings of the First Workshop on Insights from Negative Results in NLP, pages 94–99, Online, November 2020. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

We attempt to replicate a named entity recognition (NER) model implemented in a popular toolkit and discover that a critical barrier to doing so is the inconsistent evaluation of improper label sequences. We define these sequences and examine how two scorers differ in their handling of them, finding that one approach produces F1 scores approximately 0.5 points higher on the CoNLL 2003 English development and test sets. We propose best practices to increase the replicability of NER evaluations by increasing transparency regarding the handling of improper label sequences.

 
Effective Architectures for Low Resource Multilingual Named Entity Transliteration. Molly Moran and Constantine Lignos. In Alina Karakanta, Atul Kr. Ojha, Chao-Hong Liu, Jade Abbott, John Ortega, Jonathan Washington, Nathaniel Oco, Surafel Melaku Lakew, Tommi A Pirinen, Valentin Malykh, Varvara Logacheva, and Xiaobing Zhao, editors, Proceedings of the 3rd Workshop on Technologies for MT of Low Resource Languages, pages 79–86, Suzhou, China, December 2020. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

In this paper, we evaluate LSTM, biLSTM, GRU, and Transformer architectures for the task of name transliteration in a many-to-one multilingual paradigm, transliterating from 590 languages to English. We experiment with different encoder-decoder combinations and evaluate them using accuracy, character error rate, and an F-measure based on longest continuous subsequences. We find that using a Transformer for the encoder and decoder performs best, improving accuracy by over 4 points compared to previous work. We explore whether manipulating the source text by adding macrolanguage flag tokens or pre-romanizing source strings can improve performance and find that neither manipulation has a positive effect. Finally, we analyze performance differences between the LSTM and Transformer encoders when using a Transformer decoder and find that the Transformer encoder is better able to handle insertions and substitutions when transliterating.

2019

 
The Challenges of Optimizing Machine Translation for Low Resource Cross-Language Information Retrieval. Constantine Lignos, Daniel Cohen, Yen-Chieh Lien, Pratik Mehta, W. Bruce Croft, and Scott Miller. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3497–3502, Hong Kong, China, November 2019. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

When performing cross-language information retrieval (CLIR) for lower-resourced languages, a common approach is to retrieve over the output of machine translation (MT). However, there is no established guidance on how to optimize the resulting MT-IR system. In this paper, we examine the relationship between the performance of MT systems and both neural and term frequency-based IR models to identify how CLIR performance can be best predicted from MT quality. We explore performance at varying amounts of MT training data, byte pair encoding (BPE) merge operations, and across two IR collections and retrieval models. We find that the choice of IR collection can substantially affect the predictive power of MT tuning decisions and evaluation, potentially introducing dissociations between MT-only and overall CLIR performance.

 
SARAL: A Low-Resource Cross-Lingual Domain-Focused Information Retrieval System for Effective Rapid Document Triage. Elizabeth Boschee, Joel Barry, Jayadev Billa, Marjorie Freedman, Thamme Gowda, Constantine Lignos, Chester Palen-Michel, Michael Pust, Banriskhem Kayang Khonglah, Srikanth Madikeri, Jonathan May, and Scott Miller. In Marta R. Costa-jussà and Enrique Alfonseca, editors, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 19–24, Florence, Italy, July 2019. Association for Computational Linguistics. [ bib | DOI | http ]
Abstract

With the increasing democratization of electronic media, vast information resources are available in less-frequently-taught languages such as Swahili or Somali. That information, which may be crucially important and not available elsewhere, can be difficult for monolingual English speakers to effectively access. In this paper we present an end-to-end cross-lingual information retrieval (CLIR) and summarization system for low-resource languages that 1) enables English speakers to search foreign language repositories of text and audio using English queries, 2) summarizes the retrieved documents in English with respect to a particular information need, and 3) provides complete transcriptions and translations as needed. The SARAL system achieved the top end-to-end performance in the most recent IARPA MATERIAL CLIR+summarization evaluations. Our demonstration system provides end-to-end open query retrieval and summarization capability, and presents the original source text or audio, speech transcription, and machine translation, for two low resource languages.

2018

 
Combining rule-based and statistical mechanisms for low-resource named entity recognition. Ryan Gabbard, Jay DeYoung, Constantine Lignos, Marjorie Freedman, and Ralph Weischedel. Machine Translation, 32:31–43, 2018. [ bib | DOI | .pdf ]
 
The Locus of Linguistic Variation. Constantine Lignos, Laurel MacKenzie, and Meredith Tamminga, editors. John Benjamins, 2018. Originally published as a special issue of Linguistic Variation 16(2), 2016. [ bib ]
Abstract

This volume explores how the patterning of surface variation can shed light on the grammatical representation of variable phenomena. The authors explore variation in several domains, addressing intra- and inter-dialectal patterns, using diverse sources of data including corpora of naturally-occurring speech and judgment studies, and drawing on lesser-studied varieties of familiar languages, such as Northwest British Englishes and varieties of Canadian French. Ultimately, the contributions serve to expand our understanding of the nature of the mental representations and abstract processes required to support variation in language.

2017

 
Morphology and language acquisition. Constantine Lignos and Charles Yang. In Andrew Hippisley and Gregory T. Stump, editors, Cambridge Handbook of Morphology, chapter 27, pages 765–791. 2017. [ bib | .pdf ]

2016

 
The locus of linguistic variation. Constantine Lignos, Laurel MacKenzie, and Meredith Tamminga, editors. John Benjamins, December 2016. Special issue of Linguistic Variation 16(2). [ bib | DOI ]
 
Introduction: Special Issue on the Locus of Linguistic Variation. Constantine Lignos, Laurel MacKenzie, and Meredith Tamminga. Linguistic Variation, 16(2):vii--x, 2016. [ bib | DOI | http ]

2015

 
Provably correct reactive control from natural language. Autonomous Robots, 38:89–105, 2015. [ bib | DOI | .pdf ]

2014

 
Spectro-temporal correlates of lexical access during auditory lexical decision. Jonathan Brennan, Constantine Lignos, David Embick, and Timothy P.L. Roberts. Brain and Language, 133:39–46, 2014. [ bib | http ]
 
Revisiting frequency and storage in morphological processing. Constantine Lignos and Kyle Gorman. In Proceedings of the 48th Annual Meeting of the Chicago Linguistic Society, volume 48, pages 447–461, 2014. This article was submitted in November 2012, but proceedings were not officially published until 2014. [ bib ]

2013

 
You Can't Get There from Here: On Interpreting Learning Experiments. Constantine Lignos. University of Pennsylvania Working Papers in Linguistics, 19(12), 2013. [ bib | http ]
Abstract

Artificial language learning experiments provide a unique opportunity to observe learning under controlled conditions. We cannot, however, observe what learning strategy participants use; we can only carefully design the language and observe the response. This poses an inference problem that I name "the poverty of the experiment." I use computational learning models to address this inference problem, using data from an artificial grammar learning study (Saffran 2001) in which the authors conclude that participants learned hierarchical structure from distributional cues. Simulations show that that learning hierarchical structure is not required to pass the tests administered in those experiments and that a heuristic learner is the best fit for the observed human performance. Artificial language learning experiments cannot in themselves provide evidence for a particular learning strategy; they must be paired with appropriate modeling work to confirm that an implementation of a proposed learning strategy actually produces the expected results.

 
Sorry Dave, I'm Afraid I Can't Do That: Explaining Unachievable Robot Tasks Using Natural Language. Vasumathi Raman, Constantine Lignos, Cameron Finucane, Kenton CT Lee, Mitchell P Marcus, and Hadas Kress-Gazit. In Robotics: Science and Systems IX, 2013. [ bib | .pdf ]

2012

 
Infant word segmentation: An incremental, integrated model. Constantine Lignos. In Proceedings of the West Coast Conference on Formal Linguistics 30, 2012. [ bib | .pdf ]
 
Make it so: Continuous, flexible natural language interaction with an autonomous robot. Daniel J Brooks, Constantine Lignos, Cameron Finucane, Mikhail S Medvedev, Ian Perera, Vasumathi Raman, Hadas Kress-Gazit, Mitch Marcus, and Holly A Yanco. In Proceedings of the Grounding Language for Physical Systems Workshop at the Twenty-Sixth AAAI Conference on Artificial Intelligence, 2012. [ bib | http ]

2011

 
Modeling Infant Word Segmentation. Constantine Lignos. In Sharon Goldwater and Christopher Manning, editors, Proceedings of the Fifteenth Conference on Computational Natural Language Learning, pages 29–38, Portland, Oregon, USA, June 2011. Association for Computational Linguistics. [ bib | http ]

2010

 
Recession Segmentation: Simpler Online Word Segmentation Using Limited Resources. Constantine Lignos and Charles Yang. In Mirella Lapata and Anoop Sarkar, editors, Proceedings of the Fourteenth Conference on Computational Natural Language Learning, pages 88–97, Uppsala, Sweden, July 2010. Association for Computational Linguistics. [ bib | http ]
 
A Rule-Based Acquisition Model Adapted for Morphological Analysis. Constantine Lignos, Erwin Chan, Mitchell Marcus, and Charles Yang. In C. Peters, G. Di Nunzio, M. Kurimo, T. Mandl, D. Mostefa, A. Penas, and G. Roda, editors, Multilingual Information Access Evaluation I. Text Retrieval Experiments, volume 6241 of Lecture Notes in Computer Science, pages 658–665. Springer Berlin / Heidelberg, 2010. [ bib | .pdf ]
 
Investigating the Relationship Between Linguistic Representation and Computation through an Unsupervised Model of Human Morphology Learning. Erwin Chan and Constantine Lignos. Research on Language and Computation, 8:209–238, 2010. [ bib | DOI | .pdf ]
 
Learning from Unseen Data. Constantine Lignos. In Mikko Kurimo, Sami Virpioja, and Ville T. Turunen, editors, Proceedings of the Morpho Challenge 2010 Workshop, pages 35–38, Helsinki, Finland, September 2–3 2010. Aalto University School of Science and Technology. [ bib | .pdf ]
 
Evidence for a Morphological Acquisition Model from Development Data. Constantine Lignos, Erwin Chan, Charles Yang, and Mitchell P. Marcus. In Katie Franich, Kate M. Iserman, and Lauren L. Keil, editors, Proceedings of the 34th Annual Boston University Conference on Language Development. Cascadilla Press, 2010. [ bib | .pdf ]

2009

 
A Rule-Based Unsupervised Morphology Learning Framework. Constantine Lignos, Erwin Chan, Mitchell P. Marcus, and Charles Yang. In Working Notes of the 10th Workshop of the Cross-Language Evaluation Forum, CLEF 2009, Corfu, Greece, September 30--October 2 2009. [ bib | .pdf ]

2006

 
Effects of head movement on perceptions of humanoid robot behavior. Emily Wang, Constantine Lignos, Ashish Vatsal, and Brian Scassellati. In HRI '06: Proceedings of the 1st ACM SIGCHI/SIGART Conference on Human-Robot Interaction, pages 180–185, New York, NY, USA, 2006. ACM. [ bib | DOI | .pdf ]