Biotech / AI

🔥 Basecamp Research Raised $140M to Teach AI the 99% of Biology Nobody Ever Sequenced

Basecamp Research raised a $140M Series C led by S32, with Anthropic and Nvidia backing an AI trained on DNA from over a million barely-studied species to design medicine from scratch.

Basecamp Research Raised $140M to Teach AI the 99% of Biology Nobody Ever Sequenced — Tech Arcade
Photo: Sangharsh Lohakare / Unsplash

Every AI model that designs proteins today learned biology from a library that’s almost all one book. Human DNA makes up more than half of the world’s most-used public sequence archive, and five species account for roughly two-thirds of it. A London-based startup called Basecamp Research spent four years going to the other 99% — rainforest soil, volcanic vents, Antarctic seafloor — and now investors are paying up for the result.

Let’s get the number on the table first: $140 million. That’s Basecamp’s oversubscribed Series C, led by Bill Maris’s S32, with Anthropic and Nvidia’s NVentures both writing checks ($140M Series C (GEN)). The round pulls in the NATO Innovation Fund, the Rockefeller Foundation, and the UK’s Sovereign AI Fund, and brings Basecamp’s total raised to roughly $225 million (Tech.eu).

The pitch is simple to say and hard to do: if you want an AI that designs medicine, feed it the parts of life nobody has looked at yet.

đź§  Why This Matters

The last five years of “AI for biology” ran on public data — GenBank, UniProt, the Protein Data Bank. Those archives are enormous, but they’re lopsided. Basecamp’s own research pegs it bluntly: about 68% of the sequences in the Sequence Read Archive come from just five species, and humans alone are 54% (the-decoder). Train a model on that and you get a model that’s fluent in a tiny, over-studied corner of the living world.

“If you were to train an LLM only on newspaper articles from 1975, it would be a really, really bad model.” — Philip Lorenz, CTO, Basecamp Research (the-decoder)

Basecamp’s bet is that the frontier isn’t a cleverer model — it’s better fuel. Its dataset, BaseData, already holds about 15 trillion DNA “tokens” and the company says the first generation of its EDEN models trained on DNA from over one million newly sequenced species (the-decoder). That’s the moat Nvidia and Anthropic are buying into: proprietary biology you can’t scrape off the internet.

📊 Deep Dive

Basecamp runs a global sampling network across 30+ countries and all seven continents, then turns the DNA it collects into training data for EDEN, a foundation model with 28 billion parameters (SiliconANGLE). The goal isn’t to fold proteins for fun — it’s to design working biological machinery: enzymes, gene-insertion tools, and eventually cell therapies you dose in vivo, inside the body, instead of manufacturing outside it.

The early lab numbers are what got the term sheet signed:

  • Tumor clearance: AI-designed T cells cleared more than 90% of tumor cells in laboratory tests (the-decoder).
  • Gene-editing tools: Half of the large serine recombinases EDEN generated were active in human cells — a brutal bar most designed enzymes fail (GEN).
  • Hit rate: A 63.2% functional hit rate across diverse DNA prompts (GEN).
  • Antibiotics: A candidate called EDEN-7 performed comparably to last-resort antibiotics against multidrug-resistant bacteria in mice (the-decoder).
  • Synthetic ecosystems: An EDEN-generated microbiome spanning 9,067 species hit 99% biome-specific taxonomic accuracy (GEN).

For context on the trajectory: Basecamp raised a $20 million Series A in 2022 and a $60 million Series B in 2024 (Sifted). The data pile is meant to grow just as fast — the company is targeting one quadrillion tokens within 18 months through a “Trillion Gene Atlas” it’s building with Nvidia and Anthropic (the-decoder).

“We believe the future of medicine lies in reprogramming the body to repair itself. We design the models and the medicines to teach it how.” — Glen Gowers, co-founder and CEO (SiliconANGLE)

⚠️ The Catch

Read the fine print and the caution writes itself. Every headline result — the tumor clearance, the antibiotic, the recombinases — happened in a dish or in mice. Not one of these therapies has touched a human patient, and the gap between “clears 90% of tumor cells in the lab” and “approved drug” is where most biotech money goes to die. It’s usually a decade and hundreds of millions of dollars.

Basecamp also declined to disclose a valuation (GEN), which usually means the number is either very high or awkward to defend. And the collection model itself invites scrutiny: gathering genetic material across 30-plus countries raises real questions about benefit-sharing and biodiversity rights, the kind the Nagoya Protocol was written to govern. A dataset’s value as a moat is exactly what makes its sourcing worth watching.

🎯 What Happens Next

The $140 million buys two things: bigger EDEN models and a pipeline push toward the clinic, with a focus on in vivo cell therapy (Tech.eu). Watch for the first therapeutic candidate Basecamp names a target and a timeline for — that’s when the lab numbers start facing the FDA. Watch, too, for whether the data keeps compounding: hitting that quadrillion-token goal is the difference between a one-time advantage and a durable one.

And keep an eye on the backers. When a chipmaker and a frontier-model lab both fund the same biology dataset, they’re not just being nice — they want the compute demand and the model partnership that come with it.

đź§© Bigger Picture

For a decade, the scarce input in AI was compute. Then it was clean data. Basecamp is a bet on a third bottleneck: data that doesn’t exist yet — you have to go dig it out of a volcano first. If that thesis holds, the winners in AI-designed medicine won’t be whoever has the most GPUs. They’ll be whoever went to the most places nobody else bothered to sample.

It also reframes what “AI for science” means. AlphaFold predicted structures from what we already knew. Basecamp is trying to expand what we know, then let the model design against it. That’s a slower, dirtier, more expensive game — and if it works, a much larger one.

The last frontier of biology was never going to be sequenced from a laptop. It was always going to require someone willing to get soil under their fingernails — and now, $140 million to turn that dirt into medicine.


Sources