Antoine Limasset
Chargé de recherche CNRS · CRIStAL · Université de Lille
Bonsai team — Algorithms and data structures for sequence analysis
Publications · Software · Team · Funding · Teaching · Contact
About Me
I am a CNRS researcher in the Bonsai team at CRIStAL, Université de Lille. I design algorithms and data structures for large genomic and transcriptomic datasets, connecting theoretical computer science with practical bioinformatics.
My research focuses on sequence indexing, compact k-mer representations, compression, sketching and sequencing error correction. I develop open-source tools to make large-scale sequence analysis more efficient in both memory and computation.
I obtained my habilitation to direct research (HDR) at Université de Lille in September 2025.
Publications
Search the full bibliography · citations and BibTeX · Google Scholar
Download all references (BibTeX)
2026
-
ZOR Filters: Fast and Smaller Than Fuse Filters
24th International Symposium on Experimental Algorithms (SEA 2026), LIPIcs 371, 24:1–24:17. Schloss Dagstuhl – Leibniz-Zentrum für Informatik. -
Compressed inverted indexes for scalable sequence similarity
RECOMB 2026. Available text: bioRxiv preprint, first posted in 2025. -
Minimizer Density revisited: Models and Multiminimizers
RECOMB 2026. Available text: bioRxiv preprint, first posted in 2025. -
Super Bloom: Fast and precise filter for streaming k-mer queries
RECOMB-Seq 2026 ; preprint bioRxiv. -
Accelerating k-mer-based sequence filtering
Peer Community Journal, 6, e55.
2025
-
Inverted colored de Bruijn Graph for practical kmer sets storage
bioRxiv, preprint. -
Fractional hitting sets for efficient multiset sketching
Algorithms for Molecular Biology, 20 (1), 1. -
REINDEER2: Practical Abundance Index at Scale
String Processing and Information Retrieval (SPIRE 2025), Lecture Notes in Computer Science 16073, 156–171. Springer Nature. -
OReO: optimizing read order for practical compression
Bioinformatics Advances, 5 (1), vbaf128. -
K2R: Tinted de Bruijn graphs implementation for efficient read extraction from sequencing datasets
Bioinformatics Advances, 5 (1), vbaf111. -
Hyper-k-mers: Efficient Streaming k-mers Representation
Research in Computational Molecular Biology (RECOMB 2025), Lecture Notes in Computer Science 15647, 330–335. Springer Nature.
2024
-
Automated evaluation of multiple sequence alignment methods to handle third generation sequencing errors
PeerJ, 12, e17731. -
Conway–Bromage–Lyndon (CBL): an exact, dynamic representation of k-mer sets
Bioinformatics, 40 (Suppl. 1), i48–i57. -
Brisk: Exact resource-efficient dictionary for k-mers
bioRxiv, preprint. -
Evaluating k-mer Transformations for Cache Coherence and Uniformity
Zenodo, preprint.
2023
-
Locality-preserving minimal perfect hashing of k-mers
Bioinformatics, 39 (Suppl. 1), i534–i543. -
Scalable sequence database search using partitioned aggregated Bloom comb trees
Bioinformatics, 39 (Suppl. 1), i252–i259. -
Fractional Hitting Sets for Efficient and Lightweight Genomic Data Sketching
23rd International Workshop on Algorithms in Bioinformatics (WABI 2023), LIPIcs 273, 15:1–15:27. Schloss Dagstuhl – Leibniz-Zentrum für Informatik.
2022
- Toward Optimal Fingerprint Indexing for Large Scale Genomics
22nd International Workshop on Algorithms in Bioinformatics (WABI 2022), LIPIcs 242, 25:1–25:15. Schloss Dagstuhl – Leibniz-Zentrum für Informatik.
2021
-
BLight: efficient exact associative structure for k-mers
Bioinformatics, 37 (18), 2858–2865. -
Scalable long read self-correction and assembly polishing with multiple sequence alignment
Scientific Reports, 11 (1), 761.
2020
-
Toward perfect reads: self-correction of short reads via mapping on de Bruijn graphs
Bioinformatics, 36 (5), 1374–1381. -
ELECTOR: evaluator for long reads correction methods
NAR Genomics and Bioinformatics, 2 (1), lqz015. -
A resource-frugal probabilistic dictionary and applications in bioinformatics
Discrete Applied Mathematics, 274, 92–102.
2019
- Read correction for non-uniform coverages
bioRxiv, preprint.
2017
- Fast and Scalable Minimal Perfect Hashing for Massive Key Sets
16th International Symposium on Experimental Algorithms (SEA 2017), LIPIcs 75, 25:1–25:16. Schloss Dagstuhl – Leibniz-Zentrum für Informatik.
2016
-
Compacting de Bruijn graphs from sequencing data quickly and in low memory
Bioinformatics, 32 (12), i201–i208. -
Read mapping on de Bruijn graphs
BMC Bioinformatics, 17 (1), 237. -
A Resource-frugal Probabilistic Dictionary and Applications in (Meta)Genomics
Proceedings of the Prague Stringology Conference 2016, 85–98. Jan Holub and Jan Žďárek (eds.), Czech Technical University in Prague.
2015
- On the Representation of De Bruijn Graphs
Journal of Computational Biology, 22 (5), 336–352.
2014
- On the Representation of de Bruijn Graphs
Research in Computational Molecular Biology (RECOMB 2014), Lecture Notes in Computer Science 8394, 35–55. Springer.
Collaborative Publications
2022
- Critical Assessment of Metagenome Interpretation: the second round of challenges
Nature Methods, 19 (4), 429–440.
2021
-
Chromosome-level genome assembly reveals homologous chromosomes and recombination in asexual rotifer Adineta vaga
Science Advances, 7 (41), eabg4216. -
STRONG: metagenomics strain resolution on assembly graphs
Genome Biology, 22 (1), 214.
Software
Indexing, search & data structures
-
ZOR
Memory-efficient approximate membership filters.
-
SuperBloom
Fast Bloom filters for streaming k-mer queries.
-
Onika
Compressed inverted indexes for scalable sequence similarity search.
-
REINDEER2
Scalable indexing and querying of k-mer abundances.
-
K2R
Exact retrieval of reads containing specified k-mers.
-
K2Rmini
Fast k-mer-based sequence filtering.
-
KFC
K-mer counting using compact hyper-k-mer representations.
-
CBL
Exact, dynamic k-mer sets with support for set operations.
-
LPHASH
Compact locality-preserving minimal perfect hashing of k-mers.
-
PAC
Sequence database search using partitioned aggregated Bloom comb trees.
-
BLight
Efficient exact associative k-mer dictionaries.
-
Brisk
A resource-efficient exact k-mer dictionary.
-
BBHash
Fast and scalable minimal perfect hashing for massive key sets.
-
BCALM2
Compacted de Bruijn graph construction in low memory.
-
SRC
Sequence abundance estimation and read similarity search.
Compression
Sketching & similarity
-
SuperSampler
Efficient genomic sketching using fractional hitting sets.
-
NIQKI
Fingerprint indexing for large-scale genomic sketch comparisons.
Error correction, alignment & assembly
-
MSA-Limit
Evaluation of multiple sequence alignment methods for sequencing errors.
-
STRONG
Metagenomic strain resolution on assembly graphs.
-
CONSENT
Long-read self-correction and assembly polishing with multiple sequence alignment.
-
BCOOL
Short-read correction using de Bruijn graphs.
-
BGREAT
Read mapping on de Bruijn graphs.
-
BWISE
Short-read assembly for heterozygous and polyploid genomes.
-
ELECTOR
Evaluation of long-read correction methods.
-
BRRR
A long-read correction tool based on the k-mer spectrum.
Team & Supervision
Current PhD Students
-
Étienne Conchon-Kerjan — PhD director, 2026–present.
-
Yohan Hernandez-Courbevoie — PhD director, 2024–present. Indexing global transcriptomic databases.
-
Timothé Rouzé — PhD co-supervisor, 2023–present. Compression of large sequencing collections.
Research Staff
-
Lucas Robidou — Research engineer, supervisor, 2026–present. Scalable genomic sequence analysis.
-
Florian Ingels — Postdoctoral researcher, supervisor, 2025–2026. Minimizer schemes.
Former Students and Staff
-
Léa Vandamme — PhD director, 2022–2025. Indexing third-generation sequencing datasets.
-
Caleb Smith — Engineer, supervisor, 2023–2024. Compression of large sequencing collections.
-
Coralie Rohmer — PhD co-supervisor, 2019–2023. Multiple sequence alignment algorithms for third-generation sequencing.
Education
- 2025 — Habilitation à diriger des recherches (HDR), Université de Lille. Defended on 4 September 2025.
- 2017 — PhD in Computer Science, Université de Rennes 1. Novel approaches for the exploitation of high throughput sequencing data. Supervised by Pierre Peterlongo and Dominique Lavenier; defended on 12 July 2017.
- 2014 — MSc in Computer Science, École Normale Supérieure de Rennes.
- 2012 — BSc in Computer Science, École Normale Supérieure de Cachan.
Professional Experience
- 2018–present — CNRS researcher, Bonsai team, CRIStAL, Lille, France.
- 2017 — Postdoctoral researcher, Université Libre de Bruxelles, Belgium. De novo assembly of heterozygous genomes.
Grants & Funding
-
2026 — ANR PRC GRANDSMERS, principal investigator, approximately €585k. Graph-based Research on Accurate Nucleotide Data via Scalable, Multi-scale, and Efficient RepresentationS.
-
2026 — ANR PRC PRO-K-MER, member. PRObabilistic K-MERs for environmental sequence analysis.
-
2024 — MIC INSERM, principal investigator, €554k. Analyse efficace et évolutive du cancer par exploration transcriptomique avancée à grande échelle.
-
2024 — ANR Shannon x Cray, member, €500k.
-
2021 — ANR JCJC, principal investigator, €227k. Adequate graph structures for third-generation sequencing data exploration.
-
2019 — Hauts-de-France Region PhD Grant, principal investigator, €150k.
-
CDP PIE — Protein-Interaction-Evolution — Université de Lille Initiative d’Excellence. Total project funding: €1.5M over 4 years, renewable.
Teaching
- 2026 — Instructor, EMBO Practical Course on Pangenomics, Naples, Italy.
- 2026 — Training team and organiser, Scalable Genomics and Pangenomics, Wellcome Genome Campus, Hinxton, UK.
- 2020–2027 — Genome assembly course for master’s students, France.
- 2019–2026 — Genome assembly course, Evomics workshop, Czech Republic.
- 2023–2026 — Genome assembly course, CNRS Formation, France.
- 2015–2017 — Functional programming for bachelor’s students, France.
Professional Service
Conference Committees
- Program committees: RECOMB (2020–2026), ECCB/ISMB (2020–2026), SeqBim (2020–2025), ACM-BCB (2020–2024).
- Organizing committee: SPIRE (2021).
Reviewing
Nature Communications, Nature Methods, Genome Research, Genome Biology, Nucleic Acids Research, Bioinformatics, Scientific Reports, and other journals and conferences.
Thesis Committees
- Nastasija Mijovic — PhD committee, 2023–2025.
- Riku Walve — Examiner, 2022.
- Svitlana Lukicheva — PhD jury, 2021.
- Théo Lemane — PhD committee, 2020–2021.
- Nadege Guiglielmoni — PhD committee, 2019–2020.
Talks & Presentations
-
2026 — Invited keynote, JC2B — Junior Conference on Computational Biology, Gif-sur-Yvette, France.
-
2026 — RECOMB and RECOMB-Seq, Thessaloniki, Greece.
-
2024 — EMBL-EBI K-mer/sequence indexing workshop, Cambridge, UK.
-
2024 — Kmer days, Dijon, France.
-
2023 — ISMB, Lyon, France.
-
2022 — RECOMB, San Diego, US.
-
2022 — DSB, Düsseldorf, Germany.
-
2022 — TUDASTIC, Lille, France.
-
2022 — Genopim kickoff, Rennes, France.
-
2021 — Kmer days, Marville, France.
-
2019 — Biata, Saint Petersburg, Russia.
-
2018 — RECOMB, Paris, France.
Contact
Antoine Limasset
CRIStAL (UMR 9189), Université de Lille
Bâtiment ESPRIT, 59655 Villeneuve d’Ascq, France