Antoine Limasset
Chargé de recherche CNRS · CRIStAL · Université de Lille
Bonsai team — Algorithms and data structures for sequence analysis
Publications · Software · Team · Funding · Teaching · Contact
About Me
I am a CNRS researcher in the Bonsai team at CRIStAL, Université de Lille. I design algorithms and data structures for large genomic and transcriptomic datasets, connecting theoretical computer science with practical bioinformatics.
My research focuses on sequence indexing, compact k-mer representations, compression, sketching and sequencing error correction. I develop open-source tools to make large-scale sequence analysis more efficient in both memory and computation.
I obtained my habilitation to direct research (HDR) at Université de Lille in September 2025.
Publications
Search the full bibliography · citations and BibTeX · Google Scholar
Download all references (BibTeX)
2026
-
ZOR Filters: Fast and Smaller Than Fuse Filters
24th International Symposium on Experimental Algorithms (SEA 2026), LIPIcs 371, 24:1–24:17. Schloss Dagstuhl – Leibniz-Zentrum für Informatik.
DOI: 10.4230/LIPIcs.SEA.2026.24 · Software: ZOR -
Compressed inverted indexes for scalable sequence similarity
RECOMB 2026. Available text: bioRxiv preprint, first posted in 2025.
Preprint DOI: 10.1101/2025.11.21.689685 · RECOMB 2026 · Software: Onika -
Minimizer Density revisited: Models and Multiminimizers
RECOMB 2026. Available text: bioRxiv preprint, first posted in 2025.
Preprint DOI: 10.1101/2025.11.21.689688 · RECOMB 2026 -
Super Bloom: Fast and precise filter for streaming k-mer queries
RECOMB-Seq 2026 ; preprint bioRxiv.
Preprint DOI: 10.64898/2026.03.17.712354 · Software: SuperBloom -
Accelerating k-mer-based sequence filtering
Peer Community Journal, 6, e55.
DOI: 10.24072/pcjournal.735 · Software: K2Rmini
2025
-
Inverted colored de Bruijn Graph for practical kmer sets storage
bioRxiv, preprint.
Preprint DOI: 10.64898/2025.12.08.692073 · Software: KLOE -
Fractional hitting sets for efficient multiset sketching
Algorithms for Molecular Biology, 20 (1), 1.
DOI: 10.1186/s13015-024-00268-0 · Software: SuperSampler -
REINDEER2: Practical Abundance Index at Scale
String Processing and Information Retrieval (SPIRE 2025), Lecture Notes in Computer Science 16073, 156–171. Springer Nature.
DOI: 10.1007/978-3-032-05228-5_14 · Software: REINDEER2 -
OReO: optimizing read order for practical compression
Bioinformatics Advances, 5 (1), vbaf128.
DOI: 10.1093/bioadv/vbaf128 · Software: OReO -
K2R: Tinted de Bruijn graphs implementation for efficient read extraction from sequencing datasets
Bioinformatics Advances, 5 (1), vbaf111.
DOI: 10.1093/bioadv/vbaf111 · Software: K2R -
Hyper-k-mers: Efficient Streaming k-mers Representation
Research in Computational Molecular Biology (RECOMB 2025), Lecture Notes in Computer Science 15647, 330–335. Springer Nature.
DOI: 10.1007/978-3-031-90252-9_33 · Software: KFC
2024
-
Automated evaluation of multiple sequence alignment methods to handle third generation sequencing errors
PeerJ, 12, e17731.
DOI: 10.7717/peerj.17731 · Software: MSA-Limit -
Conway–Bromage–Lyndon (CBL): an exact, dynamic representation of k-mer sets
Bioinformatics, 40 (Suppl. 1), i48–i57.
DOI: 10.1093/bioinformatics/btae217 · Software: CBL -
Brisk: Exact resource-efficient dictionary for k-mers
bioRxiv, preprint.
Preprint DOI: 10.1101/2024.11.26.625346 · Software: Brisk -
Evaluating k-mer Transformations for Cache Coherence and Uniformity
Zenodo, preprint.
Preprint DOI: 10.5281/zenodo.10870713
2023
-
Locality-preserving minimal perfect hashing of k-mers
Bioinformatics, 39 (Suppl. 1), i534–i543.
DOI: 10.1093/bioinformatics/btad219 · Software: LPHASH -
Scalable sequence database search using partitioned aggregated Bloom comb trees
Bioinformatics, 39 (Suppl. 1), i252–i259.
DOI: 10.1093/bioinformatics/btad225 · Software: PAC -
Fractional Hitting Sets for Efficient and Lightweight Genomic Data Sketching
23rd International Workshop on Algorithms in Bioinformatics (WABI 2023), LIPIcs 273, 15:1–15:27. Schloss Dagstuhl – Leibniz-Zentrum für Informatik.
DOI: 10.4230/LIPIcs.WABI.2023.15
2022
- Toward Optimal Fingerprint Indexing for Large Scale Genomics
22nd International Workshop on Algorithms in Bioinformatics (WABI 2022), LIPIcs 242, 25:1–25:15. Schloss Dagstuhl – Leibniz-Zentrum für Informatik.
DOI: 10.4230/LIPIcs.WABI.2022.25 · Software: NIQKI
2021
-
BLight: efficient exact associative structure for k-mers
Bioinformatics, 37 (18), 2858–2865.
DOI: 10.1093/bioinformatics/btab217 · Software: BLight -
Scalable long read self-correction and assembly polishing with multiple sequence alignment
Scientific Reports, 11 (1), 761.
DOI: 10.1038/s41598-020-80757-5 · Software: CONSENT
2020
-
Toward perfect reads: self-correction of short reads via mapping on de Bruijn graphs
Bioinformatics, 36 (5), 1374–1381.
DOI: 10.1093/bioinformatics/btz102 · Corrigendum · Software: BCOOL -
ELECTOR: evaluator for long reads correction methods
NAR Genomics and Bioinformatics, 2 (1), lqz015.
DOI: 10.1093/nargab/lqz015 · Software: ELECTOR -
A resource-frugal probabilistic dictionary and applications in bioinformatics
Discrete Applied Mathematics, 274, 92–102.
DOI: 10.1016/j.dam.2018.03.035 · Software: SRC
2019
- Read correction for non-uniform coverages
bioRxiv, preprint.
Preprint DOI: 10.1101/673624
2017
- Fast and Scalable Minimal Perfect Hashing for Massive Key Sets
16th International Symposium on Experimental Algorithms (SEA 2017), LIPIcs 75, 25:1–25:16. Schloss Dagstuhl – Leibniz-Zentrum für Informatik.
DOI: 10.4230/LIPIcs.SEA.2017.25 · Software: BBHash
2016
-
Compacting de Bruijn graphs from sequencing data quickly and in low memory
Bioinformatics, 32 (12), i201–i208.
DOI: 10.1093/bioinformatics/btw279 · Software: BCALM2 -
Read mapping on de Bruijn graphs
BMC Bioinformatics, 17 (1), 237.
DOI: 10.1186/s12859-016-1103-9 · Software: BGREAT -
A Resource-frugal Probabilistic Dictionary and Applications in (Meta)Genomics
Proceedings of the Prague Stringology Conference 2016, 85–98. Jan Holub and Jan Žďárek (eds.), Czech Technical University in Prague.
Proceedings, pp. 85–98
2015
- On the Representation of De Bruijn Graphs
Journal of Computational Biology, 22 (5), 336–352.
DOI: 10.1089/cmb.2014.0160
2014
- On the Representation of de Bruijn Graphs
Research in Computational Molecular Biology (RECOMB 2014), Lecture Notes in Computer Science 8394, 35–55. Springer.
DOI: 10.1007/978-3-319-05269-4_4
Collaborative Publications
2022
- Critical Assessment of Metagenome Interpretation: the second round of challenges
Nature Methods, 19 (4), 429–440.
DOI: 10.1038/s41592-022-01431-4
2021
-
Chromosome-level genome assembly reveals homologous chromosomes and recombination in asexual rotifer Adineta vaga
Science Advances, 7 (41), eabg4216.
DOI: 10.1126/sciadv.abg4216 -
STRONG: metagenomics strain resolution on assembly graphs
Genome Biology, 22 (1), 214.
DOI: 10.1186/s13059-021-02419-7 · Software: STRONG
Software
Indexing, search & data structures
-
ZOR
Memory-efficient approximate membership filters.
-
SuperBloom
Fast Bloom filters for streaming k-mer queries.
-
Onika
Compressed inverted indexes for scalable sequence similarity search.
-
REINDEER2
Scalable indexing and querying of k-mer abundances.
-
K2R
Exact retrieval of reads containing specified k-mers.
-
K2Rmini
Fast k-mer-based sequence filtering.
-
KFC
K-mer counting using compact hyper-k-mer representations.
-
CBL
Exact, dynamic k-mer sets with support for set operations.
-
LPHASH
Compact locality-preserving minimal perfect hashing of k-mers.
-
PAC
Sequence database search using partitioned aggregated Bloom comb trees.
-
BLight
Efficient exact associative k-mer dictionaries.
-
Brisk
A resource-efficient exact k-mer dictionary.
-
BBHash
Fast and scalable minimal perfect hashing for massive key sets.
-
BCALM2
Compacted de Bruijn graph construction in low memory.
-
SRC
Sequence abundance estimation and read similarity search.
Compression
Sketching & similarity
-
SuperSampler
Efficient genomic sketching using fractional hitting sets.
-
NIQKI
Fingerprint indexing for large-scale genomic sketch comparisons.
Error correction, alignment & assembly
-
MSA-Limit
Evaluation of multiple sequence alignment methods for sequencing errors.
-
STRONG
Metagenomic strain resolution on assembly graphs.
-
CONSENT
Long-read self-correction and assembly polishing with multiple sequence alignment.
-
BCOOL
Short-read correction using de Bruijn graphs.
-
BGREAT
Read mapping on de Bruijn graphs.
-
BWISE
Short-read assembly for heterozygous and polyploid genomes.
-
ELECTOR
Evaluation of long-read correction methods.
-
BRRR
A long-read correction tool based on the k-mer spectrum.
Team & Supervision
Current PhD Students
-
Étienne Conchon-Kerjan — PhD director, 2026–present.
-
Yohan Hernandez-Courbevoie — PhD director, 2024–present. Indexing global transcriptomic databases.
-
Timothé Rouzé — PhD co-supervisor, 2023–present. Compression of large sequencing collections.
Research Staff
-
Lucas Robidou — Research engineer, supervisor, 2026–present. Scalable genomic sequence analysis.
-
Florian Ingels — Postdoctoral researcher, supervisor, 2025–2026. Minimizer schemes.
Former Students and Staff
-
Léa Vandamme — PhD director, 2022–2025. Indexing third-generation sequencing datasets.
-
Caleb Smith — Engineer, supervisor, 2023–2024. Compression of large sequencing collections.
-
Coralie Rohmer — PhD co-supervisor, 2019–2023. Multiple sequence alignment algorithms for third-generation sequencing.
Education
- 2025 — Habilitation à diriger des recherches (HDR), Université de Lille. Defended on 4 September 2025.
- 2017 — PhD in Computer Science, Université de Rennes 1. Novel approaches for the exploitation of high throughput sequencing data. Supervised by Pierre Peterlongo and Dominique Lavenier; defended on 12 July 2017.
- 2014 — MSc in Computer Science, École Normale Supérieure de Rennes.
- 2012 — BSc in Computer Science, École Normale Supérieure de Cachan.
Professional Experience
- 2018–present — CNRS researcher, Bonsai team, CRIStAL, Lille, France.
- 2017 — Postdoctoral researcher, Université Libre de Bruxelles, Belgium. De novo assembly of heterozygous genomes.
Grants & Funding
-
2026 — ANR PRC GRANDSMERS, principal investigator, approximately €585k. Graph-based Research on Accurate Nucleotide Data via Scalable, Multi-scale, and Efficient RepresentationS.
-
2026 — ANR PRC PRO-K-MER, member. PRObabilistic K-MERs for environmental sequence analysis.
-
2024 — MIC INSERM, principal investigator, €554k. Analyse efficace et évolutive du cancer par exploration transcriptomique avancée à grande échelle.
-
2024 — ANR Shannon x Cray, member, €500k.
-
2021 — ANR JCJC, principal investigator, €227k. Adequate graph structures for third-generation sequencing data exploration.
-
2019 — Hauts-de-France Region PhD Grant, principal investigator, €150k.
-
CDP PIE — Protein-Interaction-Evolution — Université de Lille Initiative d’Excellence. Total project funding: €1.5M over 4 years, renewable.
Teaching
- 2026 — Instructor, EMBO Practical Course on Pangenomics, Naples, Italy.
- 2026 — Training team and organiser, Scalable Genomics and Pangenomics, Wellcome Genome Campus, Hinxton, UK.
- 2020–2027 — Genome assembly course for master’s students, France.
- 2019–2026 — Genome assembly course, Evomics workshop, Czech Republic.
- 2023–2026 — Genome assembly course, CNRS Formation, France.
- 2015–2017 — Functional programming for bachelor’s students, France.
Professional Service
Conference Committees
- Program committees: RECOMB (2020–2026), ECCB/ISMB (2020–2026), SeqBim (2020–2025), ACM-BCB (2020–2024).
- Organizing committee: SPIRE (2021).
Reviewing
Nature Communications, Nature Methods, Genome Research, Genome Biology, Nucleic Acids Research, Bioinformatics, Scientific Reports, and other journals and conferences.
Thesis Committees
- Nastasija Mijovic — PhD committee, 2023–2025.
- Riku Walve — Examiner, 2022.
- Svitlana Lukicheva — PhD jury, 2021.
- Théo Lemane — PhD committee, 2020–2021.
- Nadege Guiglielmoni — PhD committee, 2019–2020.
Talks & Presentations
-
2026 — Invited keynote, JC2B — Junior Conference on Computational Biology, Gif-sur-Yvette, France.
-
2026 — RECOMB and RECOMB-Seq, Thessaloniki, Greece. RECOMB-Seq
-
2024 — EMBL-EBI K-mer/sequence indexing workshop, Cambridge, UK.
-
2024 — Kmer days, Dijon, France.
-
2023 — ISMB, Lyon, France.
-
2022 — RECOMB, San Diego, US.
-
2022 — DSB, Düsseldorf, Germany.
-
2022 — TUDASTIC, Lille, France.
-
2022 — Genopim kickoff, Rennes, France.
-
2021 — Kmer days, Marville, France.
-
2019 — Biata, Saint Petersburg, Russia. Conference abstracts DOI: 10.1186/s12859-019-3122-9
-
2018 — RECOMB, Paris, France.
Contact
Antoine Limasset
CRIStAL (UMR 9189), Université de Lille
Bâtiment ESPRIT, 59655 Villeneuve d’Ascq, France