Skip to content
View Aaramis's full-sized avatar
đź’­
Seuls les mots mal choisis sont une perte de temps
đź’­
Seuls les mots mal choisis sont une perte de temps

Block or report Aaramis

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Aaramis/README.md

Auguste Gardette — machine learning and biological sequences

explore the code  ·  say hello

Hi, I’m Auguste. I like problems where the science is hard, the data is messy, and whatever we build still has to survive contact with reality.

I move between machine learning, genomics and software engineering. Most days, that means teaching models to read biological sequences—and occasionally teaching them to admit when they do not know.

Currently: PhD researcher at IRD–UMMISCO × Sorbonne Université, exploring environmental DNA and open-world biodiversity AI within the MetaPlantCode project.

Things on my bench

CurateMake — an auditable DNA barcode curation workflow

My current open-source rabbit hole: one traceable workflow from heterogeneous public databases to harmonised taxonomy, alignments, phylogenetic checks and reports.

Now available as a bioRxiv preprint; journal submission in preparation.


DECAF — classifying contaminants in environmental DNA

An open-source framework I developed at UMMISCO to detect contaminants in environmental DNA amplicons and turn model predictions into clean, pipeline-ready sequence files.

Work in progress

HyperTax — learning taxonomic hierarchies in hyperbolic space

An active research prototype for representing taxonomic lineages with hyperbolic entailment cones. I am currently testing how well its geometry preserves hierarchy, supports unseen taxa and transfers to biological-sequence models.

The code remains private while the experiments and scientific framing are being stabilised.

Questions I keep poking

  • How should a model behave when the right answer is a species it has never seen before?
  • Can the geometry of an embedding reflect the shape of a biological hierarchy?
  • What changes when predictions meet taxonomy, ecology and environmental context?
  • How much engineering does it take before a scientific result becomes genuinely reproducible?

Side quests

BioMiMiC
Natural-molecule screening with SMILES-BERT, built in 48 hours—2nd prize at D4Gen 2024.

Protein Classifier
ProtCNN and protein language models trying to navigate 16,652 Pfam families.

Pipex
Unix pipes rebuilt in C, because sometimes the abstraction itself is the fun part.

What is usually open in my terminal?

Python · PyTorch · Hugging Face · Snakemake · Nextflow · Docker · AWS · SQL · C/C++


If you are exploring strange data, living systems or models that need a healthier relationship with uncertainty, we will probably have something to talk about.

Pinned Loading

  1. Protein_Classifier Protein_Classifier Public

    Python

  2. Hackathon_NGS_2022 Hackathon_NGS_2022 Public

    Forked from hippolyte456/Hackathon_NGS_2022

    Projet Master2 AMI2B test de reproductibilité

    HTML 1

  3. BioMiMiC BioMiMiC Public

    G4GEN 2024 Hackathon : 2nd Winner

    Python

  4. CurateMake CurateMake Public

    Python 3