Here is a number that should make every AI founder sit up: $375 million. That's the annualized revenue run rate Snorkel AI is now claiming โ€” up roughly 18x in twelve months โ€” and it comes not from a chatbot, not from an agent, not from a shiny new model, but from selling the raw material those things are trained on (TechCrunch).

On September 22, the nine-year-old startup closed a $350 million Series E that values it at $3.5 billion โ€” nearly triple the $1.3 billion it was worth in its last round back in May 2025 (Implicator.ai). The round was co-led by Insight Partners and S32, with Addition, Lightspeed, Greylock, GV and Wells Fargo along for the ride (Unite.AI).

You've heard for two years that compute is the bottleneck. The Snorkel round is a bet on the opposite idea: the models are hungry, and they've eaten the internet. What's scarce now is data good enough to make them smarter โ€” and someone has to make it.

The AI gold rush spent 2024 buying shovels. It's spending 2026 paying people who know where the gold actually is.

๐Ÿง  Why This Matters

For most of the LLM era, the recipe was crude: scrape everything, throw it at a bigger model, watch the benchmarks climb. That worked until it didn't. The public web is finite, and the easy gains from simply hoovering up more of it have flattened.

What frontier labs need now is hard data โ€” expert-written coding problems a senior engineer would sweat over for days, medical reasoning, legal analysis, the messy edge cases where models still faceplant. That's exactly what Snorkel says it sells: a mix of synthetic generation and human subject-matter experts, packaged as finished datasets rather than the labeling software it started with (TechCrunch).

"Data 2.0 is driven by the quality of more complex data โ€” and it's primarily a research and technology problem." โ€” Alex Ratner, co-founder and CEO, Snorkel AI

Translation: the next leap in AI won't come from a bigger pile of tokens. It comes from harder ones. And whoever owns the pipeline that produces them owns a piece of every model trained downstream.

๐Ÿ“Š Deep Dive

Snorkel started as a Stanford research project in 2015 and went commercial in 2019, originally selling Snorkel Flow โ€” software that let companies label data automatically instead of by hand. The 2026 story is a pivot: away from selling tools, toward selling the finished data itself through an expert network and an evaluation product. Its customers now include, in the company's own words, frontier labs, hyperscalers, "neolabs," enterprises and U.S. government agencies (Unite.AI).

And Snorkel is not alone in cashing in on the data crunch. The whole "AI data factory" category has gone vertical:

  • Snorkel AI โ€” $3.5B valuation, ~$375M revenue run rate, $350M Series E (Sept 2026).
  • Mercor โ€” hit a $10B valuation in October 2025 and has been in talks for as much as $20B, on the back of an expert-data marketplace (CNBC).
  • Handshake โ€” the former campus-recruiting app crossed $1B in ARR after pivoting hard into AI training work (Dealroom).
  • Scale AI โ€” the original data giant, whose stake Meta valued at $14.3 billion when it bought in last year (Value Add).

Put those four together and you're looking at a sector that barely existed as a headline three years ago and now carries tens of billions in combined value. The picks-and-shovels of the AI boom turned out not to be GPUs alone โ€” they're the data that makes GPUs worth turning on.

โš ๏ธ The Catch

Squint at that $375 million and you'll notice what's missing. It's a run rate, not audited annual revenue โ€” a projection of one strong period stretched across a year. Snorkel hasn't disclosed audited figures, customer concentration, gross margins, or independent tests showing its datasets beat a rival's (Implicator.ai).

Customer concentration is the quiet risk across this whole category. When your buyers are a handful of frontier labs, one canceled contract can rewrite your growth chart overnight. Mercor offers the cautionary footnote: its headline "$2B ARR" has been picked apart by analysts who put the real net figure closer to $600 million once pass-through payments to contractors are stripped out (Value Add). Data revenue can be lumpier and lower-margin than a clean SaaS number makes it look.

There's also the existential question hanging over the entire trade: what happens when synthetic data gets good enough that the models can generate their own hard problems? Snorkel is betting the human expert stays in the loop. That bet is the business.

๐ŸŽฏ What Happens Next

With $350 million in the bank, expect Snorkel to spend aggressively on its expert network โ€” the recruiters, PhDs and domain specialists who produce the high-end data โ€” and to lean harder into evaluation, the increasingly critical job of measuring whether a model is actually any good. That's a smart hedge: even in a world of infinite synthetic data, somebody still has to grade the test.

Watch the government angle, too. Snorkel naming U.S. agencies as customers points at a lucrative, sticky market for vetted, secure training data โ€” one where "we scraped it off Reddit" is not an acceptable answer.

๐Ÿงฉ Bigger Picture

The tidy story of the AI boom was always: whoever builds the smartest model wins. The messier truth of 2026 is that the model layer is commoditizing. Open weights are catching up, and the gap between the best model and the second-best keeps shrinking.

So the money is quietly migrating to the parts of the stack that don't commoditize โ€” chips at the bottom, and data at the top. A $3.5 billion valuation for a company that makes homework for machines is what it looks like when the industry admits, out loud, that the smartest model is only as good as the hardest questions you can find to teach it.

The models scraped the whole internet and still couldn't ace the exam. Snorkel's whole business is writing a harder one โ€” and charging $375 million a year to grade it.


Sources