About me

I build machine learning systems that turn text into structured knowledge: information extraction, entity linking and knowledge graphs, across general and biomedical domains. My work has produced open datasets and benchmarks that other groups build on, including DWIE, TempEL and BioDEX, alongside 24 peer-reviewed publications at ACL, NeurIPS, VLDB and CIKM.

Before returning to research I spent seven years building software in industry, including production machine learning for fraud prevention at MercadoLibre. Most recently I held a Marie Skłodowska-Curie Postdoctoral Fellowship at Aarhus University, in the Algorithms, Data and Artificial Intelligence group, working alongside the INDE Lab at the University of Amsterdam on topics related to Knowledge Engineering.

I am open to research and applied-ML roles in industry and academia.

Selected work

EMERGE — A benchmark for updating knowledge graphs as new textual knowledge appears. In submission, NeurIPS 2026. 5 ★ · code and data
DWIE — An entity-centric dataset for multi-task, document-level information extraction — now a standard benchmark in the area. Information Processing & Management, 2021. 118 citations51 ★ · code and data
BioDEX — Large-scale extraction of adverse drug events from biomedical text, for real-world pharmacovigilance. Findings of EMNLP 2023. 28 citations65 ★ · code and data
TempEL — Entity linking against a Wikipedia that changes over time: how do models cope as entities evolve and new ones appear? NeurIPS 2022, Datasets & Benchmarks. 28 citations10 ★ · code and data
Consistent document-level entity linking — Joint modelling of entity linking and coreference resolution, so a document's mentions resolve consistently. ACL 2022 (oral). 23 citations11 ★ · code and data
Influences on LLM calibration — When do large language models know what they don't know? A study of response agreement, loss functions and prompt styles. ACL 2025 (oral, SAC Highlights Award). 17 citations · code and data

Full publication list →   CV and résumé →

Technical skills

Programming: Python (10+ yrs), Java (7+ yrs), SQL, Scala, Groovy, JavaScript
Machine learning: PyTorch, Pandas, scikit-learn, NLTK, spaCy, PyTorch Geometric, NetworkX, TensorFlow
Data and retrieval: Spark, Lucene, UIMA, Tableau, R
Databases: Oracle, MySQL, Neo4j
Infrastructure: Linux, Git, Jenkins, Docker, Nginx, Flask

Contact