Luis Morgado da Costa

I am an Assistant Professor in Computational Linguistics at Vrije Universiteit (VU) Amsterdam, working within the Computational Linguistics & Text Mining Lab (CLTL). My work bridges theoretical linguistics and modern artificial intelligence, focusing on building systems and models that are interpretable, grounded, and human-centric. I have an interdisciplinary background and a strong foundation in multilingual Computational Linguistics and Natural Language Processing.

My core research interests operate at the intersection of hybrid language modeling, applied Medical and Educational NLP, and multimodal embodied AI. You can explore some of my active projects — covering grammatical error detection, automated pronunciation checking, clinical NLP extraction/classification, and improvisational comedy robots — on my Current Projects page.

I am deeply committed to open-source human-centric development of language technologies. Examples of this include my ongoing involvement with initiatives like DELPH-IN, and the Global Wordnet Association. Beyond research, I teach courses in Computational Linguistics, NLP and Conversational/Generative AI. I also supervise graduate students working on both foundational and applied topics in language technology (mostly in the educational and medical domains).

Before joining VU Amsterdam, I was a Marie Skłodowska-Curie Fellow at Palacký University Olomouc (Czechia) and completed my PhD at Nanyang Technological University (Singapore), focusing on rich computational models for grammatical error detection.

To see my academic output, please visit my Publications or check out my recent Talks.

You can find and follow me in some of the usual places:         

News & Updates


Coming Soon

Core Research Themes

My research operates at the intersection of theoretical linguistics and modern artificial intelligence. Rather than viewing language modeling solely as an engineering challenge, I focus on integrating deep linguistic knowledge with computational methods to build systems that are interpretable, grounded, and human-centric.

My work spans three interconnected themes:

  • Hybrid & Grounded Language Modeling: I explore the synthesis of theoretically-driven symbolic approaches (such as computational grammars and rich lexical semantics) with modern machine learning techniques. I am particularly interested in building small language models and understanding how their learning can be improved by grounding them in linguistic knowledge. By anchoring statistical models in explicit linguistic structures, I aim to develop NLP systems that are more robust and genuinely comprehend language.
  • Human-Centric AI in Education and Healthcare: I apply my language modeling work to high-stakes domains—primarily Educational NLP (such as Intelligent Computer-Assisted Language Learning) and Medical NLP. A guiding principle of this applied research is keeping the human in control. I focus on designing "human-in-the-loop" systems that assist, empower, and provide transparent feedback to teachers, learners, and medical professionals, rather than replacing them with opaque black boxes.
  • Computational Linguistics & Language Documentation: I leverage computational modeling to advance the science of linguistics itself. This includes using NLP to test formal linguistic hypotheses at scale, as well as applying machine learning to assist in the documentation and preservation of diverse languages. A major component of this work involves building open-source, high-quality language resources—such as wordnets and corpora—to support both technological development and linguistic research for low-resource languages.

Recently, my research has expanded into the field of social and communicative robots, particularly working with the Leolani platform. I am fascinated by the multimodal dimensions of human-robot interaction—especially how to effectively ground conversational speech, text, and visual signals to facilitate natural and meaningful communication. Aligning with my broader focus on human-centric AI, I am actively exploring how these communicative robots can be deployed as helpful, interactive assistants within both educational and healthcare settings.

Current & Archived Projects


Below is an overview of my active and archived research projects. But since academic work naturally evolves across different funding streams and phases, the line between current and past is fluid. Projects are shown as archived when there hasn't been any active development for a while or when there is no expectation that it will resume development in the near future.


Active Projects


The AI-based Prediction of Recovery of Functioning (A-PROOF) project is an interdisciplinary Medical NLP initiative. Its primary goal is to develop and apply advanced natural language processing techniques to automatically extract critical information from free-text clinical notes within Dutch Electronic Health Records (EHRs).

To standardize this extraction, the project leverages the World Health Organization's International Classification of Functioning, Disability and Health (ICF) framework. By building deep language models capable of parsing complex medical Dutch, the system can identify, categorize, and track various domains of a patient's functioning over time. This structured data is then utilized to predict trajectories for the recovery of functioning, providing transparent and valuable decision-support for rehabilitation and healthcare professionals.

Aligning with my core research theme of Human-Centric AI in Healthcare, this project operates through a close collaboration between computational linguists at Vrije Universiteit Amsterdam, medical experts at Amsterdam UMC, and rehabilitation centers. It aims to empower medical staff with data-driven insights rather than replacing their clinical judgment.

This project is a collaboration with Laura Rupp, directly tied to her MOOC, English Pronunciation in a Global World. The primary focus is on developing automatic assessment tools for English pronunciation using modern NLP and speech processing technology. Crucially, rather than enforcing a strict native-speaker standard, this work evaluates pronunciation through the practical lens of global intelligibility—focusing on how well learners can be understood in international contexts.

Between 2025 and 2026, the foundational work was supported by funding from the Network Institute at VU Amsterdam under the project EN-SPEAK (English Speech and Pronunciation Enhancement AI Kit). The ongoing technical and linguistic development remains highly collaborative and is currently being driven by student researchers through dedicated internships and MA theses.

Leo the Robot
Leo playing with comedians!

This project is developed in collaboration with Joanna Sio, a linguist and professional comedian. It centers around the creation of an embodied agent named Leoric — which stands for Leolani-based Embodied Robot for Improvised Comedy (though its friends just call it Leo). The system uses the Leolani platform as its foundational architecture.

Leo is developed to be an aspirant comedian, and programmed to understand and utilize the principles of improv comedy to actively play improv games with human partners. Beyond the entertainment value, this project serves as a rich research vehicle. It deeply investigates aspects of embodied AI and human-robot interaction (HRI) in a dynamic, live performance setting, while also contributing broadly to the study of computational and AI-generated humor.

Publications & Related Work:

The Cantonese Wordnet, developed in collaboration with Joanna Sio (Palacký University), aims to provide a linguistically rich, open-source lexicon of Hong Kong Cantonese. Beyond its function as a database, the project is driven by a broader mission to promote and ensure Cantonese continues to thrive in the digital age. It serves as a foundational resource for Cantonese lexical semantics, enabling new documentation and computational research.

Since its inception, the project has expanded significantly. It now includes the companion Open Cantonese Sense-Tagged Corpus, providing rich contextual data for computational linguistics. Most recently, we have integrated audio recordings into the architecture, taking steps toward a "talking Wordnet" to properly capture the spoken nuances of the language.

A major and growing focus of the project is on education, with a special emphasis on supporting learners of Cantonese. Current development aims at creating a CEFR-aligned graded lexicon for Cantonse, along with digital platforms supported by our wordnet to make learning Cantonese more accessible, structured, and technologically enhanced.

Publications & Related Work:

This project focuses on the continuous development of the Mandarin Resource Grammar (MRG), an open-source, deep computational grammar for Mandarin Chinese. Developed as part of the DELPH-IN (Deep Linguistic Processing with HPSG Initiative) consortium, the grammar is grounded in the Head-driven Phrase Structure Grammar (HPSG) framework. The overarching goal is to provide a theoretically rigorous precision grammar capable of deep semantic parsing and generation for Mandarin Chinese. The project's source code are actively maintained on GitHub.

The theoretical and computational foundations of this work originate from ZHONG, a broad-coverage open-source HPSG grammar for Mandarin Chinese that also served as one of the primary foci of my PhD dissertation. While ZHONG established the core linguistic architecture, the MRG represents the next evolutionary step in this research. It actively refines the theoretical and computational implementations with a specific focus on Mandarin Chinese (instead of being a meta-Chinese grammar), supporting a better development cycle and clearer application targets.

A major application of this precision grammar lies in the educational domain, specifically targeting Intelligent Computer-Assisted Language Learning (ICALL). Between 2021 and 2023, this effort was supported by a Marie Skłodowska-Curie Action fellowship funded by the European Commission, under the project name Chinese Intelligent Language Learning (CHILL). CHILL expanded the grammar's capabilities by designing and implementing mal-rules—specialized grammatical rules designed to intentionally parse and diagnose common errors made by learners of Mandarin Chinese with a focus on the Mandarin Chinese NP structure.

Publications & Related Work:

The project explores the application of deep linguistic parsing models to build robust, intelligent Technology Enhanced Language Learning tools for English. At its core, the system utilizes the English Resource Grammar (ERG) (source code here), a broad-coverage precision grammar. To perform grammatical error detection, iTELL uses mal-rules provided by the ERG — the exact same technology I am concurrently developing for Mandarin Chinese — allowing the parser to actively identify and diagnose common learner mistakes.

iTELL started as one of the primary foci of my PhD studies. Early development of the project focused heavily on supporting English academic writing. This effort was anchored by the creation of two major data resources: the NTU Corpus of Learner's English (NTUCLE) and the Tembusu Treebank. Developed in collaboration with the Language and Communication Centre at NTU. These learner corpora contain writing samples from engineering students, richly annotated for grammatical and stylistic errors to test and refine the parser's capabilities.

Over time, the project's scope expanded beyond formal academic writing to make language practice more fun and accessible. This led to the development of CALLIG (Computer Assisted Language Learning using Improvisation Games), which links to my other projects involving computational humour, and which successfully adapted the underlying iTELL parsing architecture into interactive, gamified language learning environments.

Publications & Related Work:
  • Luís Morgado da Costa, Francis Bond, and Roger V. P. Winder. 2022. The Tembusu Treebank: An English Learner Treebank. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 4817–4826, Marseille, France. European Language Resources Association.
  • Luís Morgado da Costa, Roger V P Winder, Shu Yun Li, Benedict Christopher Lin Tzer Liang, Joseph Mackinnon, and Francis Bond. 2020. Automated Writing Support Using Deep Linguistic Parsers. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 369–377, Marseille, France. European Language Resources Association.
  • Luís Morgado da Costa and Joanna Ut-Seong Sio. 2020. CALLIG: Computer Assisted Language Learning using Improvisation Games. In Workshop on Games and Natural Language Processing, pages 49–58, Marseille, France. European Language Resources Association.
  • Roger Vivek Placidus Winder, Joseph MacKinnon, Shu Yun Li, Benedict Christopher Tzer Liang Lin, Carmel Lee Hah Heah, Luís Morgado da Costa, Takayuki Kuribayashi, and Francis Bond. 2017. NTUCLE: Developing a Corpus of Learner English to Provide Writing Support for Engineering Students. In Proceedings of the 4th Workshop on Natural Language Processing Techniques for Educational Applications (NLPTEA 2017), pages 1–11, Taipei, Taiwan. Asian Federation of Natural Language Processing.

The Abui Wordnet is a project born from a collaboration with František Kratochvíl. This project focuses on bridging traditional field linguistics methodology with computational lexical semantics. Its primary goal is to build and extend a high-quality wordnet for Abui, an extremely low-resource Papuan language spoken in eastern Indonesia. This project serves as a model for integrating endangered and under-documented languages into the global digital infrastructure.

To overcome the data scarcity typical of low-resource settings, the project creatively leverages existing language documentation resources. Specifically, it bootstraps the wordnet using data collected directly in the field, transforming rich Toolbox dictionaries into structured semantic networks. We further expand the resource by aligning and integrating the outputs of SIL's Rapid Word Collection (RWC) workshops, demonstrating how community-driven field linguistics workflows can directly feed into modern NLP resources.

Publications & Related Work:
  • Luis Morgado da Costa, František Kratochvíl, George Saad, Benidiktus Delpada, Daniel Simon Lanma, Francis Bond, Natálie Wolfová, and A.L. Blake. 2023. Linking SIL Semantic Domains to Wordnet and Expanding the Abui Wordnet through Rapid Word Collection Methodology. In Proceedings of the 12th Global Wordnet Conference, pages 315–324, University of the Basque Country, Donostia - San Sebastian, Basque Country. Global Wordnet Association.
  • Frantisek Kratochvil and Luís Morgado da Costa. 2022. Abui Wordnet: Using a Toolbox Dictionary to develop a wordnet for a low-resource language. In Proceedings of the First Workshop on NLP applications to field linguistics, pages 54–63, Gyeongju, Republic of Korea. International Conference on Computational Linguistics.
  • Luis Morgado da Costa, Francis Bond and František Kratochvíl. 2016. Linking and Disambiguating Swadesh Lists: Expanding the Open Multilingual Wordnet Using Open Language Resources. Proceedings of GLOBALEX 2016 Lexicographic Resources for Human Language Technology, 10th edition of the International Conference on Language Resources and Evaluation (LREC 2016). Portorož, Slovenia.

Archived Projects


The Coptic Wordnet was an international Digital Humanities collaboration primarily with Laura Slaughter and So Miyagawa, among others. Contextualized within the broader scope of Coptic Studies, a major innovation of this project was its experimental approach to fully automated wordnet building. The team leveraged multiple dictionaries across both ancient and modern languages to bootstrap the lexical database. The project's data is maintained open-source on GitHub.

The primary goal behind developing this automated resource was to tackle the complex challenge of text reuse in ancient manuscripts. By applying the semantic relationships mapped within the Coptic Wordnet, we were able to conduct early experiments in intertextuality studies, specifically focusing on the automatic detection of shared texts and citations within Coptic monastic writings.

Publications & Related Work:

Between 2015 and 2023, I was actively involved in the development and expansion of the Chinese Open Wordnet (COW), a resource originally developed in parallel with the NTU Multilingual Corpus. During this time, I coordinated multiple phases of the project, which included training students to sense-tag the corpus, expanding the wordnet with classifiers, chengyu, and exclamatives, translating Princeton WordNet definitions into Mandarin Chinese, and evaluating its potential as an educational tool.

While standalone development on COW is currently mostly dormant, the project laid important groundwork for computational Mandarin semantics. In the future, this resource is likely to be revived and extended by directly linking its lexical database with the deep linguistic representations of the Mandarin Resource Grammar (MRG).

Publications & Related Work:
  • Ut Seong Sio and Luís Morgado da Costa. 2022. Multilingual Reference Annotation: A Case between English and Mandarin Chinese. In Proceedings of the 18th Joint ACL - ISO Workshop on Interoperable Semantic Annotation within LREC2022, pages 86–94, Marseille, France. European Language Resources Association.
  • Francis Bond, Andrew Devadason, Melissa Rui Lin Teo, and Luís Morgado da Costa. 2021. Teaching Through Tagging — Interactive Lexical Semantics. In Proceedings of the 11th Global Wordnet Conference, pages 273–283, University of South Africa (UNISA). Global Wordnet Association.
  • Francis Bond, Tomoko Ohkuma, Luis Morgado da Costa, Yasuhide Miura, Rachel Chen, Takayuki Kuribayashi and Wenjie Wang. 2016. A Multilingual Sentiment Corpus for Chinese, English and Japanese. Proceedings of Emotion and Sentiment Analysis Workshop, 10th edition of the International Conference on Language Resources and Evaluation (LREC 2016). Portorož, Slovenia
  • Luis Morgado Da Costa, Francis Bond, and Helena Gao. 2016. Mapping and Generating Classifiers using an Open Chinese Ontology. In Proceedings of the 8th Global WordNet Conference (GWC), pages 249–256, Bucharest, Romania. Global Wordnet Association.
  • Francis Bond, Luís Morgado da Costa, and Tuấn Anh Lê. 2015. IMI — A Multilingual Semantic Annotation Environment. In Proceedings of ACL-IJCNLP 2015 System Demonstrations, pages 7–12, Beijing, China. Association for Computational Linguistics and The Asian Federation of Natural Language Processing.
  • Luís Morgado da Costa and Francis Bond. 2015. OMWEdit - The Integrated Open Multilingual Wordnet Editing System. In Proceedings of ACL-IJCNLP 2015 System Demonstrations, pages 73–78, Beijing, China. Association for Computational Linguistics and The Asian Federation of Natural Language Processing.

During my time at Nanyang Technological University and in the years following, I was a primary contributor to the Open Multilingual Wordnet (OMW) and its supporting infrastructure — a project coordinated by Francis Bond. My contributions spanned both front and back-end development for this rich multilingual semantic resource. This included redesigning database schemas, creating the OMWEdit web service, and developing IMI, a comprehensive Multilingual Semantic Annotation Environment used for tagging the NTU Multilingual Corpus.

On a global scale, these technical efforts supported the broader mission of the Global Wordnet Association (GWA). I helped develop the online governance and machinery for the Collaborative Interlingual Index (CILI) and the Global Wordnet Grid (GWG). This large-scale collaborative effort enables independent wordnet projects worldwide to enrich and link their concepts into a single, unified interlingual index.

While I am no longer involved in the active maintenance or core development of these centralized repositories, I remain a dedicated member of the GWA. The computational and theoretical foundations established during this period are still deeply linked to my current research, heavily informing my active development of individual open-source wordnets for languages such as Mandarin Chinese, Cantonese, Abui, Coptic, and Kristang.

Publications & Related Work:
  • Francis Bond, Michael Wayne Goodman, Ewa Rudnicka, Luis Morgado da Costa, Alexandre Rademaker, and John P. McCrae. 2023. Documenting the Open Multilingual Wordnet. In Proceedings of the 12th Global Wordnet Conference, pages 150–157, University of the Basque Country, Donostia - San Sebastian, Basque Country. Global Wordnet Association.
  • Francis Bond, Andrew Devadason, Melissa Rui Lin Teo, and Luís Morgado da Costa. 2021. Teaching Through Tagging — Interactive Lexical Semantics. In Proceedings of the 11th Global Wordnet Conference, pages 273–283, University of South Africa (UNISA). Global Wordnet Association.
  • Francis Bond, Andrew Devadason, Melissa Rui Lin Teo, and Luís Morgado da Costa. 2021. Teaching Through Tagging — Interactive Lexical Semantics. In Proceedings of the 11th Global Wordnet Conference, pages 273–283, University of South Africa (UNISA). Global Wordnet Association.
  • John P. McCrae, Michael Wayne Goodman, Francis Bond, Alexandre Rademaker, Ewa Rudnicka, and Luís Morgado Da Costa. 2021. The GlobalWordNet Formats: Updates for 2020. In Proceedings of the 11th Global Wordnet Conference, pages 91–99, University of South Africa (UNISA). Global Wordnet Association.
  • Francis Bond, Hiroki Nomoto, Luís Morgado da Costa, and Arthur Bond. 2020. Linking the TUFS Basic Vocabulary to the Open Multilingual Wordnet. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 3181–3188, Marseille, France. European Language Resources Association.
  • Francis Bond, Luis Morgado da Costa, Michael Wayne Goodman, John P. McCrae, and Ahti Lohk. 2020. Some Issues with Building a Multilingual Wordnet. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 3189–3197, Marseille, France. European Language Resources Association.
  • Laura Slaughter, Wenjie Wang, Luis Morgado Da Costa, and Francis Bond. 2018. Enchancing the Collaborative Interlingual Index for Digital Humanities: Cross-linguistic Analysis in the Domain of Theology. In Proceedings of the 9th Global Wordnet Conference, pages 341–346, Nanyang Technological University (NTU), Singapore. Global Wordnet Association.
  • Luis Morgado Da Costa and Francis Bond. 2016. Wow! What a Useful Extension! Introducing Non-Referential Concepts to Wordnet. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), pages 4323–4328, Portorož, Slovenia. European Language Resources Association (ELRA).
  • Luis Morgado da Costa, Francis Bond and František Kratochvíl. 2016. Linking and Disambiguating Swadesh Lists: Expanding the Open Multilingual Wordnet Using Open Language Resources. Proceedings of GLOBALEX 2016 Lexicographic Resources for Human Language Technology, 10th edition of the International Conference on Language Resources and Evaluation (LREC 2016). Portorož, Slovenia.
  • Francis Bond, Tomoko Ohkuma, Luis Morgado da Costa, Yasuhide Miura, Rachel Chen, Takayuki Kuribayashi and Wenjie Wang. 2016. A Multilingual Sentiment Corpus for Chinese, English and Japanese. Proceedings of Emotion and Sentiment Analysis Workshop, 10th edition of the International Conference on Language Resources and Evaluation (LREC 2016). Portorož, Slovenia
  • Francis Bond, Luís Morgado da Costa, and Tuấn Anh Lê. 2015. IMI — A Multilingual Semantic Annotation Environment. In Proceedings of ACL-IJCNLP 2015 System Demonstrations, pages 7–12, Beijing, China. Association for Computational Linguistics and The Asian Federation of Natural Language Processing.
  • Luís Morgado da Costa and Francis Bond. 2015. OMWEdit - The Integrated Open Multilingual Wordnet Editing System. In Proceedings of ACL-IJCNLP 2015 System Demonstrations, pages 73–78, Beijing, China. Association for Computational Linguistics and The Asian Federation of Natural Language Processing.

The Open Kristang Wordnet and its companion online dictionary, Pinchah Kristang, were developed within the context of Kodrah Kristang (“Awaken, Kristang”), a grassroots community initiative to revitalize Kristang—a critically endangered language in Singapore and Malaysia. I volunteered with the initiative between 2017 and 2022, providing technical guidance and support for digital efforts aimed at the language's education and maintenance.

While active development on these specific tools concluded in 2023 and the project is currently dormant, the linguistic data and framework produced during this time remain valuable. The project may eventually be revived and expanded as part of a broader, comparative Portuguese-based creole digital lexicon.

Publications & Related Work:
  • Luís Morgado da Costa. 2020. Pinchah Kristang: A Dictionary of Kristang. In Proceedings of the 2020 Globalex Workshop on Linked Lexicography, pages 37–44, Marseille, France. European Language Resources Association.

Within the QTLeap project, at NLX-Group – University of Lisbon, I was mainly responsible for the maintenance and development of Language Resources for Deep Machine Translation, namely a Portuguese Deep Parallel TreeBanking using the LxGram.

My main responsibilities at the Centro Virtual Camões, Camões, I.P. – Institute for Cooperation and Language included managing and producing online content concerning Portuguese culture and language teaching worldwide; maintaining the Camões Digital Library; as well as giving support to the e-learning center and Portuguese language certification division.

As a member of the Centre for Comparative Studies – Faculty of Letters, University of Lisbon, I was mainly responsible for the conceptualization and development of an online database destined to collect and analyse primary written sources of Portuguese Orientalism.

Courses


Currently Teaching

Previously Taught

Student Supervision


I actively supervise BA, MA, and PhD students in areas broadly related to my core research themes. I welcome motivated students interested in working on projects involving hybrid language modeling, Educational NLP (such as grammatical error detection and assessment of pronunciation quality), Medical NLP (such as information extraction from clinical records), and embodied AI or conversational robots.

Many of my students conduct research that directly contributes to active interdisciplinary collaborations, such as evaluating global English intelligibility models for EN-SPEAK, performing clinical NLP extraction for A-PROOF, or developing multimodal grounding systems for communicative robotics.

Unless otherwise advertised in dedicated channels (e.g., social media, VU's official channel, linguistlist, corpora list), our university does not provide general funding for PhD students. If you would like to do a PhD with me, please ensure you have your own source of funding before applying.


Current Students

Previous Students (Completed)

2026
2025
2024
2023

Publications


Invited Talks & Conference Presentations


Datasets


Coming Soon

Software & Tools


Coming Soon

Academic Memberships

I am an active member of several international research communities and consortia dedicated to computational linguistics, open-source language resources, and artificial intelligence.


Reviewing & Committees

I am actively involved in the academic community, serving on organizing committees and regularly participating in the peer-review process for a variety of journals, conferences, and workshops in Computational Linguistics, NLP, and Language Technology.

Conference Organizing Committees
Reviewing & Program/Scientific Committees

Luis Morgado da Costa, Computational Linguistics and Text Mining Lab, VU Amsterdam
Last modified: September 2026