Skip to content
View HenryNdubuaku's full-sized avatar
  • Cactus Compute
  • London
  • 19:25 (UTC +01:00)

Block or report HenryNdubuaku

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
HenryNdubuaku/README.md

Henry Ndubuaku

LinkedIn Twitter Email Spotify

I could train a 1B-A200m model on an iPhone 17 Pro at ~650 tokens/sec. It will take 360 days on 20B tokens of training data and use 156KW of electricity which cost $51.

The phone will fry of course, so I wrote algorithms to run inference on your phone rather. We named it after a plant that survives in resource-constrained environments, the Cactus.

cactus can run similar model on your Grandma’s Pixel 6a at 80 tokens/second while only draining 10% battery per hour of continuous inference and using 250MB RAM only.

I also built an agentic foundation model for tiny smart devices that runs 1.2k toks/sec on Cactus, its called Needle.

Needle has no FFN, just 26m params of pure attention, yet matches 10-25x bigger models. It can be a better Siri/Alexa that powers tiny smart devices.

We raised some money from YCombinator, Oxford's Seed Fund, FCVC (Slack, Coinbase, GitLab, Instacart etc.), and 6 smaller funds like Transpose (run by Garry Tan's brother), fellow YC founders, and 62 tech CTOs/VP/Dir both via syndicate and directly at Google DeepMind etc.

Cactus now powers cool products you've probably heard of...I think. 6 exceptionally gifted "Cactus Jacks" from UCLA, Nokia, Google, Stanford, Oxford, have joined us!

Same destination, just a different route!

Publications

Parameter-Efficient Transformer Embedding via Functional Factorization (ICML 2026)
HiDRA: A Blazing Fast LM-Head Replacement (ICLR 2026)
Depth Over Specialization in Small Multimodal Transformers (ICLR 2026)
Just Enough Learning: GRPO-Guided Controllers for Hyperparameter Sweeps (ICLR 2026)
TACE: Token-Aware Chunked Encoding For Realtime Speech Models (ICLR 2026)
CLAWS: Calibration-Aware Activation Sparsity for Instruction-Tuned LLMs (ICML 2026)

Fun Facts

  • After CUDARepo, Nvidia reached out, I did 7 technical rounds, got a verbal offer, back-and-forth over YOE/pay, then I got YC.
  • Did MSc at QMUL, just to work with Prof Matt Purver (Ex-Stanford Researcher on CALO), did my project/thesis with his team.
  • Did BEng under Prof Onyema Uzoamaka (Rumoured first Nigerian CS grad from MIT), he taught computing archs off-head!
  • Biggest career miss was a PhD Studentship at Meta FAIR.

Personal Life

Profile Nigerian-British, born Jan 1996, 185cm, 83kg
Hobbies Calisthenics, UFC, chess, music
Philosophy Humanist Christian, unpolitical

Speaking at UCLA

The Cactus Jacks

Team dinner

I'm fit too

Cactus Jacks at YC

Cactus Pod at YC HQ

Music Profile

Expressive Rap
Alternative/Folk
Soul/Jazz
Oldies
Genre-Blending Urban
Dark Pop

Pinned Loading

  1. maths-cs-ai-compendium maths-cs-ai-compendium Public

    Become a cracked AI/ML Research Engineer

    TypeScript 7k 837

  2. cactus-compute/cactus cactus-compute/cactus Public

    Tiny AI for tiny devices

    C++ 5.5k 447

  3. cactus-compute/needle cactus-compute/needle Public

    26m agentic model for tiny devices

    Python 3.2k 244

  4. vector vector Public

    Programming language for ML on XLA

    Rust 9 2

  5. nanodl nanodl Public

    JAX library for training sub-4B foundation models for edge

    Python 306 13

  6. cuda-tutorials cuda-tutorials Public

    Comprehensive CUDA tutorials for Maths & ML with examples

    Cuda 238 9