I am an Assistant Professor in Computational Linguistics at Vrije Universiteit (VU) Amsterdam, working within the Computational Linguistics & Text Mining Lab (CLTL). My work bridges theoretical linguistics and modern artificial intelligence, focusing on building systems and models that are interpretable, grounded, and human-centric. I have an interdisciplinary background and a strong foundation in multilingual Computational Linguistics and Natural Language Processing.
My core research interests operate at the intersection of hybrid language modeling, applied Medical and Educational NLP, and multimodal embodied AI. You can explore some of my active projects — covering grammatical error detection, automated pronunciation checking, clinical NLP extraction/classification, and improvisational comedy robots — on my Current Projects page.
I am deeply committed to open-source human-centric development of language technologies. Examples of this include my ongoing involvement with initiatives like DELPH-IN, and the Global Wordnet Association. Beyond research, I teach courses in Computational Linguistics, NLP and Conversational/Generative AI. I also supervise graduate students working on both foundational and applied topics in language technology (mostly in the educational and medical domains).
Before joining VU Amsterdam, I was a Marie Skłodowska-Curie Fellow at Palacký University Olomouc (Czechia) and completed my PhD at Nanyang Technological University (Singapore), focusing on rich computational models for grammatical error detection.
To see my academic output, please visit my Publications or check out my recent Talks.
News & Updates
Coming Soon
Core Research Themes
My research operates at the intersection of theoretical linguistics and modern artificial intelligence. Rather than viewing language modeling solely as an engineering challenge, I focus on integrating deep linguistic knowledge with computational methods to build systems that are interpretable, grounded, and human-centric.
My work spans three interconnected themes:
- Hybrid & Grounded Language Modeling: I explore the synthesis of theoretically-driven symbolic approaches (such as computational grammars and rich lexical semantics) with modern machine learning techniques. I am particularly interested in building small language models and understanding how their learning can be improved by grounding them in linguistic knowledge. By anchoring statistical models in explicit linguistic structures, I aim to develop NLP systems that are more robust and genuinely comprehend language.
- Human-Centric AI in Education and Healthcare: I apply my language modeling work to high-stakes domains—primarily Educational NLP (such as Intelligent Computer-Assisted Language Learning) and Medical NLP. A guiding principle of this applied research is keeping the human in control. I focus on designing "human-in-the-loop" systems that assist, empower, and provide transparent feedback to teachers, learners, and medical professionals, rather than replacing them with opaque black boxes.
- Computational Linguistics & Language Documentation: I leverage computational modeling to advance the science of linguistics itself. This includes using NLP to test formal linguistic hypotheses at scale, as well as applying machine learning to assist in the documentation and preservation of diverse languages. A major component of this work involves building open-source, high-quality language resources—such as wordnets and corpora—to support both technological development and linguistic research for low-resource languages.
Recently, my research has expanded into the field of social and communicative robots, particularly working with the Leolani platform. I am fascinated by the multimodal dimensions of human-robot interaction—especially how to effectively ground conversational speech, text, and visual signals to facilitate natural and meaningful communication. Aligning with my broader focus on human-centric AI, I am actively exploring how these communicative robots can be deployed as helpful, interactive assistants within both educational and healthcare settings.
Current & Archived Projects
Below is an overview of my active and archived research projects. But since academic work naturally evolves across different funding streams and phases, the line between current and past is fluid. Projects are shown as archived when there hasn't been any active development for a while or when there is no expectation that it will resume development in the near future.
Active Projects
An interdisciplinary Medical NLP project applying natural language processing to Dutch electronic health records and conversations with a clinical context to automatically extract patient functional status and predict recovery.
The AI-based Prediction of Recovery of Functioning (A-PROOF) project is an interdisciplinary Medical NLP initiative. Its primary goal is to develop and apply advanced natural language processing techniques to automatically extract critical information from free-text clinical notes within Dutch Electronic Health Records (EHRs).
To standardize this extraction, the project leverages the World Health Organization's International Classification of Functioning, Disability and Health (ICF) framework. By building deep language models capable of parsing complex medical Dutch, the system can identify, categorize, and track various domains of a patient's functioning over time. This structured data is then utilized to predict trajectories for the recovery of functioning, providing transparent and valuable decision-support for rehabilitation and healthcare professionals.
Aligning with my core research theme of Human-Centric AI in Healthcare, this project operates through a close collaboration between computational linguists at Vrije Universiteit Amsterdam, medical experts at Amsterdam UMC, and rehabilitation centers. It aims to empower medical staff with data-driven insights rather than replacing their clinical judgment.
An interdisciplinary project focusing on the AI-driven assessment of English pronunciation and global intelligibility, heavily supported by student research.
This project is a collaboration with Laura Rupp, directly tied to her MOOC, English Pronunciation in a Global World. The primary focus is on developing automatic assessment tools for English pronunciation using modern NLP and speech processing technology. Crucially, rather than enforcing a strict native-speaker standard, this work evaluates pronunciation through the practical lens of global intelligibility—focusing on how well learners can be understood in international contexts.
Between 2025 and 2026, the foundational work was supported by funding from the Network Institute at VU Amsterdam under the project EN-SPEAK (English Speech and Pronunciation Enhancement AI Kit). The ongoing technical and linguistic development remains highly collaborative and is currently being driven by student researchers through dedicated internships and MA theses.
An embodied agent designed to perform improv comedy with humans, exploring the intersection of human-robot interaction and AI-generated humor.
This project is developed in collaboration with Joanna Sio, a linguist and professional comedian. It centers around the creation of an embodied agent named Leoric — which stands for Leolani-based Embodied Robot for Improvised Comedy (though its friends just call it Leo). The system uses the Leolani platform as its foundational architecture.
Leo is developed to be an aspirant comedian, and programmed to understand and utilize the principles of improv comedy to actively play improv games with human partners. Beyond the entertainment value, this project serves as a rich research vehicle. It deeply investigates aspects of embodied AI and human-robot interaction (HRI) in a dynamic, live performance setting, while also contributing broadly to the study of computational and AI-generated humor.
Publications & Related Work:- Sio, Joanna Ut-Seong and Morgado da Costa, Luis. 2024. Humor as a mind-engineering tool in the digital age: the case of stand-up comedy. A Routledge Handbook of Language and Mind Engineering. Routledge
An open-source, foundational resource for Cantonese lexical semantics, featuring a companion corpus and audio recordings, aimed at helping the language thrive in the digital age and supporting its learners worldwide.
The Cantonese Wordnet, developed in collaboration with Joanna Sio (Palacký University), aims to provide a linguistically rich, open-source lexicon of Hong Kong Cantonese. Beyond its function as a database, the project is driven by a broader mission to promote and ensure Cantonese continues to thrive in the digital age. It serves as a foundational resource for Cantonese lexical semantics, enabling new documentation and computational research.
Since its inception, the project has expanded significantly. It now includes the companion Open Cantonese Sense-Tagged Corpus, providing rich contextual data for computational linguistics. Most recently, we have integrated audio recordings into the architecture, taking steps toward a "talking Wordnet" to properly capture the spoken nuances of the language.
A major and growing focus of the project is on education, with a special emphasis on supporting learners of Cantonese. Current development aims at creating a CEFR-aligned graded lexicon for Cantonse, along with digital platforms supported by our wordnet to make learning Cantonese more accessible, structured, and technologically enhanced.
Publications & Related Work:- Sio, Joanna Ut-Seong and Morgado da Costa, Luis. 2026. The CantoneseWordnet: an Open Lexicographic Database of Cantonese. A Routledge Handbook of Cantonese Linguistics. Routledge.
- Morgado da Costa, Luis and Sio, Joanna Ut-Seong and Chin, Andy and Lam, Zoe Wai-Man and Pai, Raymond. 2026. Promoting Cantonese in and through the DigitalWorld. A Routledge Handbook of Cantonese Linguistics. Routledge.
- Joanna Ut-Seong Sio, Luis Morgado Da Costa, Francis Bond, and Kamila Liedermannova. 2025. Can you hear me now? Towards talking Wordnets: A Cantonese Case Study. In Proceedings of the 13th Global Wordnet Conference, pages 243–248, Pavia, Italy. Global Wordnet Association.
- Joanna Sio and Luis Morgado Da Costa. 2023. The Open Cantonese Sense-Tagged Corpus. In Proceedings of the 12th Global Wordnet Conference, pages 263–268, University of the Basque Country, Donostia - San Sebastian, Basque Country. Global Wordnet Association.
- Ut Seong Sio and Luís Morgado da Costa. 2022. Enriching Linguistic Representation in the Cantonese Wordnet and Building the New Cantonese Wordnet Corpus. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 70–78, Marseille, France. European Language Resources Association.
- Joanna Ut-Seong Sio and Luis Morgado Da Costa. 2019. Building the Cantonese Wordnet. In Proceedings of the 10th Global Wordnet Conference, pages 206–215, Wroclaw, Poland. Global Wordnet Association.
The development of a theoretically grounded precision grammar and parser for Mandarin Chinese, extended with specialized mal-rules to detect and diagnose common grammatical errors made by language learners.
This project focuses on the continuous development of the Mandarin Resource Grammar (MRG), an open-source, deep computational grammar for Mandarin Chinese. Developed as part of the DELPH-IN (Deep Linguistic Processing with HPSG Initiative) consortium, the grammar is grounded in the Head-driven Phrase Structure Grammar (HPSG) framework. The overarching goal is to provide a theoretically rigorous precision grammar capable of deep semantic parsing and generation for Mandarin Chinese. The project's source code are actively maintained on GitHub.
The theoretical and computational foundations of this work originate from ZHONG, a broad-coverage open-source HPSG grammar for Mandarin Chinese that also served as one of the primary foci of my PhD dissertation. While ZHONG established the core linguistic architecture, the MRG represents the next evolutionary step in this research. It actively refines the theoretical and computational implementations with a specific focus on Mandarin Chinese (instead of being a meta-Chinese grammar), supporting a better development cycle and clearer application targets.
A major application of this precision grammar lies in the educational domain, specifically targeting Intelligent Computer-Assisted Language Learning (ICALL). Between 2021 and 2023, this effort was supported by a Marie Skłodowska-Curie Action fellowship funded by the European Commission, under the project name Chinese Intelligent Language Learning (CHILL). CHILL expanded the grammar's capabilities by designing and implementing mal-rules—specialized grammatical rules designed to intentionally parse and diagnose common errors made by learners of Mandarin Chinese with a focus on the Mandarin Chinese NP structure.
Publications & Related Work:- Luis Morgado da Costa and Francis Bond. Grammatical Error Detection Using HPSG Grammars: Diagnosing Common Mandarin Chinese Grammatical Errors. Proceedings of the 29th International Conference on Head-Driven Phrase Structure Grammar (HPSG 2022). Online, Japan
- Luís Morgado da Costa, Francis Bond, and Roger V. P. Winder. 2022. The Tembusu Treebank: An English Learner Treebank. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 4817–4826, Marseille, France. European Language Resources Association.
- Ut Seong Sio and Luís Morgado da Costa. 2022. Multilingual Reference Annotation: A Case between English and Mandarin Chinese. In Proceedings of the 18th Joint ACL - ISO Workshop on Interoperable Semantic Annotation within LREC2022, pages 86–94, Marseille, France. European Language Resources Association.
- Luís Morgado da Costa. 2021. Using Rich Models of Language in Grammatical Error Detection. Ph.D. thesis, Nanyang Technological University.
- Luis Morgado da Costa, Francis Bond, and Xiaoling He. 2016. Syntactic Well-Formedness Diagnosis and Error-Based Coaching in Computer Assisted Language Learning using Machine Translation. In Proceedings of the 3rd Workshop on Natural Language Processing Techniques for Educational Applications (NLPTEA2016), pages 107–116, Osaka, Japan. The COLING 2016 Organizing Committee.
A comprehensive framework for English language learning and error detection, combining deep linguistic parsing, learner corpora, and gamification.
The project explores the application of deep linguistic parsing models to build robust, intelligent Technology Enhanced Language Learning tools for English. At its core, the system utilizes the English Resource Grammar (ERG) (source code here), a broad-coverage precision grammar. To perform grammatical error detection, iTELL uses mal-rules provided by the ERG — the exact same technology I am concurrently developing for Mandarin Chinese — allowing the parser to actively identify and diagnose common learner mistakes.
iTELL started as one of the primary foci of my PhD studies. Early development of the project focused heavily on supporting English academic writing. This effort was anchored by the creation of two major data resources: the NTU Corpus of Learner's English (NTUCLE) and the Tembusu Treebank. Developed in collaboration with the Language and Communication Centre at NTU. These learner corpora contain writing samples from engineering students, richly annotated for grammatical and stylistic errors to test and refine the parser's capabilities.
Over time, the project's scope expanded beyond formal academic writing to make language practice more fun and accessible. This led to the development of CALLIG (Computer Assisted Language Learning using Improvisation Games), which links to my other projects involving computational humour, and which successfully adapted the underlying iTELL parsing architecture into interactive, gamified language learning environments.
Publications & Related Work:- Luís Morgado da Costa, Francis Bond, and Roger V. P. Winder. 2022. The Tembusu Treebank: An English Learner Treebank. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 4817–4826, Marseille, France. European Language Resources Association.
- Luís Morgado da Costa, Roger V P Winder, Shu Yun Li, Benedict Christopher Lin Tzer Liang, Joseph Mackinnon, and Francis Bond. 2020. Automated Writing Support Using Deep Linguistic Parsers. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 369–377, Marseille, France. European Language Resources Association.
- Luís Morgado da Costa and Joanna Ut-Seong Sio. 2020. CALLIG: Computer Assisted Language Learning using Improvisation Games. In Workshop on Games and Natural Language Processing, pages 49–58, Marseille, France. European Language Resources Association.
- Roger Vivek Placidus Winder, Joseph MacKinnon, Shu Yun Li, Benedict Christopher Tzer Liang Lin, Carmel Lee Hah Heah, Luís Morgado da Costa, Takayuki Kuribayashi, and Francis Bond. 2017. NTUCLE: Developing a Corpus of Learner English to Provide Writing Support for Engineering Students. In Proceedings of the 4th Workshop on Natural Language Processing Techniques for Educational Applications (NLPTEA 2017), pages 1–11, Taipei, Taiwan. Asian Federation of Natural Language Processing.
The development of a rich computational lexical resource by bridging field linguistics methods and language documentation resources and computational semantics to build a wordnet for an extremely low-resource Papuan language.
The Abui Wordnet is a project born from a collaboration with František Kratochvíl. This project focuses on bridging traditional field linguistics methodology with computational lexical semantics. Its primary goal is to build and extend a high-quality wordnet for Abui, an extremely low-resource Papuan language spoken in eastern Indonesia. This project serves as a model for integrating endangered and under-documented languages into the global digital infrastructure.
To overcome the data scarcity typical of low-resource settings, the project creatively leverages existing language documentation resources. Specifically, it bootstraps the wordnet using data collected directly in the field, transforming rich Toolbox dictionaries into structured semantic networks. We further expand the resource by aligning and integrating the outputs of SIL's Rapid Word Collection (RWC) workshops, demonstrating how community-driven field linguistics workflows can directly feed into modern NLP resources.
Publications & Related Work:- Luis Morgado da Costa, František Kratochvíl, George Saad, Benidiktus Delpada, Daniel Simon Lanma, Francis Bond, Natálie Wolfová, and A.L. Blake. 2023. Linking SIL Semantic Domains to Wordnet and Expanding the Abui Wordnet through Rapid Word Collection Methodology. In Proceedings of the 12th Global Wordnet Conference, pages 315–324, University of the Basque Country, Donostia - San Sebastian, Basque Country. Global Wordnet Association.
- Frantisek Kratochvil and Luís Morgado da Costa. 2022. Abui Wordnet: Using a Toolbox Dictionary to develop a wordnet for a low-resource language. In Proceedings of the First Workshop on NLP applications to field linguistics, pages 54–63, Gyeongju, Republic of Korea. International Conference on Computational Linguistics.
- Luis Morgado da Costa, Francis Bond and František Kratochvíl. 2016. Linking and Disambiguating Swadesh Lists: Expanding the Open Multilingual Wordnet Using Open Language Resources. Proceedings of GLOBALEX 2016 Lexicographic Resources for Human Language Technology, 10th edition of the International Conference on Language Resources and Evaluation (LREC 2016). Portorož, Slovenia.
Archived Projects
An experimental Digital Humanities project focusing on the fully automated construction of a wordnet to facilitate the automatic detection of text reuse in ancient manuscripts.
The Coptic Wordnet was an international Digital Humanities collaboration primarily with Laura Slaughter and So Miyagawa, among others. Contextualized within the broader scope of Coptic Studies, a major innovation of this project was its experimental approach to fully automated wordnet building. The team leveraged multiple dictionaries across both ancient and modern languages to bootstrap the lexical database. The project's data is maintained open-source on GitHub.
The primary goal behind developing this automated resource was to tackle the complex challenge of text reuse in ancient manuscripts. By applying the semantic relationships mapped within the Coptic Wordnet, we were able to conduct early experiments in intertextuality studies, specifically focusing on the automatic detection of shared texts and citations within Coptic monastic writings.
Publications & Related Work:- So Miyagawa, Luis Morgado da Costa, Laura Slaughter, and Heike Behlmer. 2025. Automatic Detection of Coptic Text Reuse: Applying Coptic Wordnet to Intertextuality Studies in Selected Coptic Monastic Writings. In Proceedings of the 13th Global Wordnet Conference, pages 179–184, Pavia, Italy. Global Wordnet Association.
- Laura Slaughter, Luis Morgado Da Costa, So Miyagawa, Marco Büchler, Amir Zeldes, and Heike Behlmer. 2019. The Making of Coptic Wordnet. In Proceedings of the 10th Global Wordnet Conference, pages 166–175, Wroclaw, Poland. Global Wordnet Association.
Past contributions to the expansion and sense-tagging of a Mandarin Chinese wordnet, with future plans for integration with the Mandarin Resource Grammar.
Between 2015 and 2023, I was actively involved in the development and expansion of the Chinese Open Wordnet (COW), a resource originally developed in parallel with the NTU Multilingual Corpus. During this time, I coordinated multiple phases of the project, which included training students to sense-tag the corpus, expanding the wordnet with classifiers, chengyu, and exclamatives, translating Princeton WordNet definitions into Mandarin Chinese, and evaluating its potential as an educational tool.
While standalone development on COW is currently mostly dormant, the project laid important groundwork for computational Mandarin semantics. In the future, this resource is likely to be revived and extended by directly linking its lexical database with the deep linguistic representations of the Mandarin Resource Grammar (MRG).
Publications & Related Work:- Ut Seong Sio and Luís Morgado da Costa. 2022. Multilingual Reference Annotation: A Case between English and Mandarin Chinese. In Proceedings of the 18th Joint ACL - ISO Workshop on Interoperable Semantic Annotation within LREC2022, pages 86–94, Marseille, France. European Language Resources Association.
- Francis Bond, Andrew Devadason, Melissa Rui Lin Teo, and Luís Morgado da Costa. 2021. Teaching Through Tagging — Interactive Lexical Semantics. In Proceedings of the 11th Global Wordnet Conference, pages 273–283, University of South Africa (UNISA). Global Wordnet Association.
- Francis Bond, Tomoko Ohkuma, Luis Morgado da Costa, Yasuhide Miura, Rachel Chen, Takayuki Kuribayashi and Wenjie Wang. 2016. A Multilingual Sentiment Corpus for Chinese, English and Japanese. Proceedings of Emotion and Sentiment Analysis Workshop, 10th edition of the International Conference on Language Resources and Evaluation (LREC 2016). Portorož, Slovenia
- Luis Morgado Da Costa, Francis Bond, and Helena Gao. 2016. Mapping and Generating Classifiers using an Open Chinese Ontology. In Proceedings of the 8th Global WordNet Conference (GWC), pages 249–256, Bucharest, Romania. Global Wordnet Association.
- Francis Bond, Luís Morgado da Costa, and Tuấn Anh Lê. 2015. IMI — A Multilingual Semantic Annotation Environment. In Proceedings of ACL-IJCNLP 2015 System Demonstrations, pages 7–12, Beijing, China. Association for Computational Linguistics and The Asian Federation of Natural Language Processing.
- Luís Morgado da Costa and Francis Bond. 2015. OMWEdit - The Integrated Open Multilingual Wordnet Editing System. In Proceedings of ACL-IJCNLP 2015 System Demonstrations, pages 73–78, Beijing, China. Association for Computational Linguistics and The Asian Federation of Natural Language Processing.
Foundational contributions to global, open-source lexical semantics infrastructure of the Global Wordnet Association, supporting the collaborative development, maintenance and interlingual linking of many wordnets worldwide.
During my time at Nanyang Technological University and in the years following, I was a primary contributor to the Open Multilingual Wordnet (OMW) and its supporting infrastructure — a project coordinated by Francis Bond. My contributions spanned both front and back-end development for this rich multilingual semantic resource. This included redesigning database schemas, creating the OMWEdit web service, and developing IMI, a comprehensive Multilingual Semantic Annotation Environment used for tagging the NTU Multilingual Corpus.
On a global scale, these technical efforts supported the broader mission of the Global Wordnet Association (GWA). I helped develop the online governance and machinery for the Collaborative Interlingual Index (CILI) and the Global Wordnet Grid (GWG). This large-scale collaborative effort enables independent wordnet projects worldwide to enrich and link their concepts into a single, unified interlingual index.
While I am no longer involved in the active maintenance or core development of these centralized repositories, I remain a dedicated member of the GWA. The computational and theoretical foundations established during this period are still deeply linked to my current research, heavily informing my active development of individual open-source wordnets for languages such as Mandarin Chinese, Cantonese, Abui, Coptic, and Kristang.
Publications & Related Work:- Francis Bond, Michael Wayne Goodman, Ewa Rudnicka, Luis Morgado da Costa, Alexandre Rademaker, and John P. McCrae. 2023. Documenting the Open Multilingual Wordnet. In Proceedings of the 12th Global Wordnet Conference, pages 150–157, University of the Basque Country, Donostia - San Sebastian, Basque Country. Global Wordnet Association.
- Francis Bond, Andrew Devadason, Melissa Rui Lin Teo, and Luís Morgado da Costa. 2021. Teaching Through Tagging — Interactive Lexical Semantics. In Proceedings of the 11th Global Wordnet Conference, pages 273–283, University of South Africa (UNISA). Global Wordnet Association.
- Francis Bond, Andrew Devadason, Melissa Rui Lin Teo, and Luís Morgado da Costa. 2021. Teaching Through Tagging — Interactive Lexical Semantics. In Proceedings of the 11th Global Wordnet Conference, pages 273–283, University of South Africa (UNISA). Global Wordnet Association.
- John P. McCrae, Michael Wayne Goodman, Francis Bond, Alexandre Rademaker, Ewa Rudnicka, and Luís Morgado Da Costa. 2021. The GlobalWordNet Formats: Updates for 2020. In Proceedings of the 11th Global Wordnet Conference, pages 91–99, University of South Africa (UNISA). Global Wordnet Association.
- Francis Bond, Hiroki Nomoto, Luís Morgado da Costa, and Arthur Bond. 2020. Linking the TUFS Basic Vocabulary to the Open Multilingual Wordnet. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 3181–3188, Marseille, France. European Language Resources Association.
- Francis Bond, Luis Morgado da Costa, Michael Wayne Goodman, John P. McCrae, and Ahti Lohk. 2020. Some Issues with Building a Multilingual Wordnet. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 3189–3197, Marseille, France. European Language Resources Association.
- Laura Slaughter, Wenjie Wang, Luis Morgado Da Costa, and Francis Bond. 2018. Enchancing the Collaborative Interlingual Index for Digital Humanities: Cross-linguistic Analysis in the Domain of Theology. In Proceedings of the 9th Global Wordnet Conference, pages 341–346, Nanyang Technological University (NTU), Singapore. Global Wordnet Association.
- Luis Morgado Da Costa and Francis Bond. 2016. Wow! What a Useful Extension! Introducing Non-Referential Concepts to Wordnet. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), pages 4323–4328, Portorož, Slovenia. European Language Resources Association (ELRA).
- Luis Morgado da Costa, Francis Bond and František Kratochvíl. 2016. Linking and Disambiguating Swadesh Lists: Expanding the Open Multilingual Wordnet Using Open Language Resources. Proceedings of GLOBALEX 2016 Lexicographic Resources for Human Language Technology, 10th edition of the International Conference on Language Resources and Evaluation (LREC 2016). Portorož, Slovenia.
- Francis Bond, Tomoko Ohkuma, Luis Morgado da Costa, Yasuhide Miura, Rachel Chen, Takayuki Kuribayashi and Wenjie Wang. 2016. A Multilingual Sentiment Corpus for Chinese, English and Japanese. Proceedings of Emotion and Sentiment Analysis Workshop, 10th edition of the International Conference on Language Resources and Evaluation (LREC 2016). Portorož, Slovenia
- Francis Bond, Luís Morgado da Costa, and Tuấn Anh Lê. 2015. IMI — A Multilingual Semantic Annotation Environment. In Proceedings of ACL-IJCNLP 2015 System Demonstrations, pages 7–12, Beijing, China. Association for Computational Linguistics and The Asian Federation of Natural Language Processing.
- Luís Morgado da Costa and Francis Bond. 2015. OMWEdit - The Integrated Open Multilingual Wordnet Editing System. In Proceedings of ACL-IJCNLP 2015 System Demonstrations, pages 73–78, Beijing, China. Association for Computational Linguistics and The Asian Federation of Natural Language Processing.
Past development of digital resources for the revitalization of Kristang, with future plans to expand into a comparative Portuguese-based creole lexicon.
The Open Kristang Wordnet and its companion online dictionary, Pinchah Kristang, were developed within the context of Kodrah Kristang (“Awaken, Kristang”), a grassroots community initiative to revitalize Kristang—a critically endangered language in Singapore and Malaysia. I volunteered with the initiative between 2017 and 2022, providing technical guidance and support for digital efforts aimed at the language's education and maintenance.
While active development on these specific tools concluded in 2023 and the project is currently dormant, the linguistic data and framework produced during this time remain valuable. The project may eventually be revived and expanded as part of a broader, comparative Portuguese-based creole digital lexicon.
Publications & Related Work:- Luís Morgado da Costa. 2020. Pinchah Kristang: A Dictionary of Kristang. In Proceedings of the 2020 Globalex Workshop on Linked Lexicography, pages 37–44, Marseille, France. European Language Resources Association.
Within the QTLeap project, at NLX-Group – University of Lisbon, I was mainly responsible for the maintenance and development of Language Resources for Deep Machine Translation, namely a Portuguese Deep Parallel TreeBanking using the LxGram.
My main responsibilities at the Centro Virtual Camões, Camões, I.P. – Institute for Cooperation and Language included managing and producing online content concerning Portuguese culture and language teaching worldwide; maintaining the Camões Digital Library; as well as giving support to the e-learning center and Portuguese language certification division.
As a member of the Centre for Comparative Studies – Faculty of Letters, University of Lisbon, I was mainly responsible for the conceptualization and development of an online database destined to collect and analyse primary written sources of Portuguese Orientalism.
Courses
Currently Teaching
-
Conversational AI and Robots [2026—] (Graduate, co-coordinating with Piek Vossen and Bram Willemsen)
A general introduction to embodied, multimodal conversational AI — with a focus on CLTL's own robot platform, Leolani. -
Conversational AI [2026—] (Undergraduate, co-coordinating with Filip Ilievski and Lea Krause)
A general introduction of modern conversational AI, progressing from foundational theories to state-of-the-art LLM architectures and applications. -
Applied Text Mining Methods [2024—] (Graduate, co-coordinating it with Isa Maks)
Hands-on practice developing modern NLP pipelines to applied language tasks. -
Academic English Grammar [2025—] (Undergraduate)
Provides advanced instruction on English grammatical structures for academic writing and analysis. -
Programming in Python for Text Analysis [2023—] (Graduate)
A foundational programming for graduate-level non-technical students specifically tailored for processing and analyzing textual data. -
Introduction to Python for Humanities and Social Sciences [2023—] (Undergraduate)
A foundational programming course designed for non-technical students for Digital Humanities. -
Thesis Seminar [2024—] (Graduate, involved as one of many lecturers)
A graduate course students guiding students through diverse topics in research design and writing process.
Previously Taught
-
LLMs in Linguistics [2025] (Graduate)
Advanced LOT School course, exploring the theoretical and practical implications and role of Large Language Models for linguistic research. -
Advanced NLP [2023—2025] (Graduate)
Advanced course focusing on evaluation of state-of-the-art architectures in Natural Language Processing, with a focus on interpretability and black-box testing. -
NLP Research Seminar [2024] (Graduate)
A yearly-themed graduate seminar covering breakthroughs, experimental designs, and limitations in a specific subfield of Computational Linguistics. -
Experiments in NLP [2023] (Graduate)
An advanced hands-on course focusing on experimental design, evaluation metrics, and reproducibility in NLP tasks. -
Topics in Chinese Syntax and Semantics [2022] (Undergraduate/Graduate)
An advanced course looking into complex grammatical structures and meaning construction in Mandarin Chinese from a theoretical and computational perspective. -
Chinese Semantics and Lexicology [2022] (Undergraduate)
An introductory course of Chinese semantics and lexicology, with a focus on computational lexical semantic resources. -
Corpus Linguistics [2020] (Undergraduate/Graduate, involved as Teaching Assistant/Tutor)
A general introduction to corpus linguistics — covering compilation, annotation, and empirical analysis of textual corpora. -
Semantics and Pragmatics [2018, 2020] (Undergraduate, involved as Teaching Assistant/Tutor)
A core introduction to semantics and pragmatics to linguistic students. -
Detecting Meaning with Sherlock Holmes [2018, 2019] (Undergraduate, involved as Teaching Assistant/Tutor)
An introductory to core semantic concepts to a broad interdisciplinary audience themed around Sherlock Holmes stories.
Student Supervision
I actively supervise BA, MA, and PhD students in areas broadly related to my core research themes. I welcome motivated students interested in working on projects involving hybrid language modeling, Educational NLP (such as grammatical error detection and assessment of pronunciation quality), Medical NLP (such as information extraction from clinical records), and embodied AI or conversational robots.
Many of my students conduct research that directly contributes to active interdisciplinary collaborations, such as evaluating global English intelligibility models for EN-SPEAK, performing clinical NLP extraction for A-PROOF, or developing multimodal grounding systems for communicative robotics.
Unless otherwise advertised in dedicated channels (e.g., social media, VU's official channel, linguistlist, corpora list), our university does not provide general funding for PhD students. If you would like to do a PhD with me, please ensure you have your own source of funding before applying.
Current Students
- Ellie Smith (PhD) – Topic: Methodologically Sound Applications of NLP for Social Sciences and Humanities (co-supervised with Antske Fokkens)
Previous Students (Completed)
2026
- Ella Heinävaara (MA) – Evaluating Speech-to-IPA Models for Intelligibility Classification of L2 English Pronunciation
- Hannah Altgassen (MA) – Grammatical Gender Prediction of English Loanwords in German Using Transformer-Based Language Models (co-supervised with Antske Fokkens)
- Marit Veerle Rozendaal (MA) – Modelling Patient Functioning Over Time: Automated ICF Categorisation and Its Representation in an Event-Centred Knowledge Graph (co-supervised with Piek Vossen)
- Masoumeh Javadi Zarnaghi (MA) – Evaluating Multimodal LLMs for Word-Level Intelligibility Classification and Phonetic Explanation
2025
- Farnaz Bani Fatemi (ReMA) – Grammaticality and LLMs: Evaluating the Potential of BabyLMs for Grammatical Error Detection in NLP
- Elisabetta Dentico (MA) – Grammatical Error Detection in L2 English and Italian: How Multilingual LLMs Handle Ambiguity in Learner Errors
- Wayne Kuan (MA) – Evaluating the Impact of Continuous Pre-Training on ASR Models for Word-Level English Pronunciation Intelligibility
- Szabolcs Pál (ReMA) – Investigation of scalable audio based speaker identification in the context of Communicative Robots (co-supervised with Piek Vossen)
- Xin Chen (MA) – Generating Follow-up Questions in Health Conversations Using Fine-tuned Language Models (co-supervised with Piek Vossen)
2024
- Furong Zou (MA) – Exploring An Existing ASR Model for a Binary Classification of Intelligibility on MOOC English Speech Data (co-supervises with Pia Sommerauer)
- Long Ma (MA) – Chinese Healthcare Named Entity Recognition (CHNER) Using BiLSTM-CRF Classifiers
- Nynke van't Hof (MSci) – Automating Performance Status Annotation in Oncology Patient Records Using Llama-3 (co-supervised with Piek Vossen)
2023
- Adam Tucker (MA) – An investigation of complex word identification (CWI) systems for English (co-supervised with Hennie Van der Vliet)
- Ajda Efendi (MA) – Document Classification on EQF levels with Multilingual datasets in English (co-supervised with Hennie Van der Vliet)
- Irma Tuinenga (MA) – Words Made Easy: a Comparative Study of Methods for English Lexical Simplification (co-supervised with Hennie Van der Vliet)
- Siti Nurhalima (MA) – Enhancing Wordnet Bahasa through Multilingual Sense Intersection (co-supervised with Hennie Van der Vliet)
- Swarupa Hardikar (MA) – Exploring Open-source Generative Models for Lexical Simplification through Prompt Learning (co-supervised with Hennie Van der Vliet)
Publications
- Loading publications...
Invited Talks & Conference Presentations
- Loading talks...
Datasets
Coming Soon
Software & Tools
Coming Soon
Academic Memberships
I am an active member of several international research communities and consortia dedicated to computational linguistics, open-source language resources, and artificial intelligence.
-
Computational Linguistics & Text Mining Lab (CLTL)
As an Assistant Professor at Vrije Universiteit (VU) Amsterdam, I am part of the core research team advancing the intersection of linguistics, computational linguistics and (embodied) AI. -
Network Institute (VU Amsterdam)
I am affiliated with the Network Institute, an interdisciplinary research hub at VU Amsterdam exploring the digital society. This membership directly supports my collaborative, human-centric work at the intersection of AI, digital humanities, and educational technology. Within the Network Institute, I co-coordinate the Academy Assistant program with my colleague Pia Sommarauer. Through this program, we help our university connect disciplines by funding bright young master students to conduct interdisciplinary research in small research projects. -
DELPH-IN (Deep Linguistic Processing with HPSG Initiative)
I am currently a member of DELPH-IN's standing committee. I share this consortium's communal commitment to the open-source development of NLP tools for high quality, linguistically motivated syntactic and semantic parsing. I am also the main developer and maintainer of the Mandarin Resource Grammar, and the DELPH-IN demo: Grammarium. -
Global Wordnet Association (GWA)
I contribute to open-source research on computational lexical semantics. My involvement includes supporting the collaborative development of multilingual wordnets and facilitating the governance of the Collaborative Interlingual Index (CILI). -
European Association of Chinese Linguistics (EACL)
Aligning with my core research in Chinese NLP and general Mandarin Chinese linguistics, I am an active member of EACL, which promotes the study and research of Chinese linguistics across Europe.
Reviewing & Committees
I am actively involved in the academic community, serving on organizing committees and regularly participating in the peer-review process for a variety of journals, conferences, and workshops in Computational Linguistics, NLP, and Language Technology.
Conference Organizing Committees
- 2027 – Global Wordnet Conference (GWC) (Local Organizing Committee)
- 2025 – The 21st DELPH-IN Summit (Local Organizing Committee)
- 2018 – GLOBALEX: Lexicography & WordNets (Organizing Committee)
- 2018 – The 9th Global WordNet Conference (GWC) (Local Organizing Committee)
- 2015 – The 11th DELPH-IN Summit (Local Organizing Committee)
- 2015 – The 22nd International Conference on Head-Driven Phrase Structure Grammar (HPSG) (Local Organizing Committee)
- 2014 – The 1st Workshop/Hackathon for the Wordnet Bahasa (Local Organizing Committee)
Reviewing & Program/Scientific Committees
- ACL Rolling Review (ARR)
- Language Resources and Evaluation Conference (LREC / LREC-COLING)
- Global WordNet Conference (GWC)
- Head-Driven Phrase Structure Grammar Conference (HPSG)
- Computer Speech & Language (Elsevier)
- Northern European Journal of Language Technology (NEJLT)
- Linguamática
- International Conference on Advances in Semantic Processing (SEMAPRO)
- Symposium on Languages, Applications and Technologies (SLATE)
- European Association for Chinese Studies (EACS)
Last modified: September 2026