explore the code  · say hello
Hi, I’m Auguste. I like problems where the science is hard, the data is messy, and whatever we build still has to survive contact with reality.
I move between machine learning, genomics and software engineering. Most days, that means teaching models to read biological sequences—and occasionally teaching them to admit when they do not know.
Currently: PhD researcher at IRD–UMMISCO × Sorbonne Université, exploring environmental DNA and open-world biodiversity AI within the MetaPlantCode project.
My current open-source rabbit hole: one traceable workflow from heterogeneous public databases to harmonised taxonomy, alignments, phylogenetic checks and reports.
Now available as a bioRxiv preprint; journal submission in preparation.
An open-source framework I developed at UMMISCO to detect contaminants in environmental DNA amplicons and turn model predictions into clean, pipeline-ready sequence files.
An active research prototype for representing taxonomic lineages with hyperbolic entailment cones. I am currently testing how well its geometry preserves hierarchy, supports unseen taxa and transfers to biological-sequence models.
The code remains private while the experiments and scientific framing are being stabilised.
- How should a model behave when the right answer is a species it has never seen before?
- Can the geometry of an embedding reflect the shape of a biological hierarchy?
- What changes when predictions meet taxonomy, ecology and environmental context?
- How much engineering does it take before a scientific result becomes genuinely reproducible?
BioMiMiC
Natural-molecule screening with SMILES-BERT, built in 48 hours—2nd prize at D4Gen 2024.
Protein Classifier
ProtCNN and protein language models trying to navigate 16,652 Pfam families.
Pipex
Unix pipes rebuilt in C, because sometimes the abstraction itself is the fun part.
What is usually open in my terminal?
Python · PyTorch · Hugging Face · Snakemake · Nextflow · Docker · AWS · SQL · C/C++
If you are exploring strange data, living systems or models that need a healthier relationship with uncertainty, we will probably have something to talk about.
