Six years ago, Snorkel AI was a Stanford lab project with a clever idea about labeling data. This week it became a $3.5 billion company selling the one raw material the entire AI industry is running short on.

Let's get the number on the table first: $350 million. That's the Series E that Snorkel just closed, led by Insight Partners and S32, with Alphabet's GV, Greylock, Lightspeed, Addition, and Wells Fargo along for the ride (TechCrunch). The round values the company at $3.5 billion โ€” up from $1.3 billion in May 2025, when it raised $100 million (Reuters).

Here's the figure that actually explains the valuation, though. Snorkel's annualized revenue run rate went from roughly $20 million to $375 million in about twelve months โ€” an 18x jump โ€” after it stopped mainly selling software and started selling data as a service (SiliconANGLE).

The thesis is simple: the scarcest resource in AI right now isn't compute or talent โ€” it's high-quality, expert-made training data, and Snorkel figured out how to manufacture it at scale.

๐Ÿง  Why This Matters

Every frontier model you've heard of was trained on a diet of internet text. That diet is largely eaten. The public web has been scraped, tokenized, and squeezed, and the labs have discovered that the next increment of capability doesn't come from more data โ€” it comes from harder data. Data that teaches a model to reason through a legal contract, debug a codebase, or work through a graduate-level chemistry problem.

That kind of data doesn't exist lying around. It has to be made, by people who actually know the subject. Snorkel's pitch is that it can produce it โ€” expert-generated examples, evaluation rubrics for grading model outputs, and the sandboxed "environments" where models practice tasks through reinforcement learning.

"Data is becoming more rare, more specialized, more difficult to find." โ€” Andy Harrison, partner at S32 (Reuters)

When the money agrees with the founders, valuations move fast. Snorkel's nearly tripled in about sixteen months.

๐Ÿ“Š Deep Dive

Snorkel started in 2019 as a spinout from the Stanford AI Lab, built around "programmatic labeling" โ€” the idea that you could tag training data with code instead of armies of human annotators. For years it sold that as software. The problem: labs didn't want a tool, they wanted the finished data. So last September, Snorkel flipped the model and started delivering the data itself. Revenue did this:

  • Snorkel AI: ~$375M annualized run rate, up ~18x in a year; $3.5B valuation
  • Mercor: roughly $2B in gross annualized revenue (TechCrunch)
  • Handshake: crossed the $1B annualized revenue milestone
  • Micro1: around $500M gross run rate
  • Scale AI: the category giant โ€” Meta paid $14.3 billion for a 49% stake in June 2025 (Reuters)

CEO and co-founder Alex Ratner put the growth plainly:

"Since launching our new data-as-a-service offering nearly a year ago, we've grown over 18 times, and this week crossed an annualized revenue run rate of $375 million." โ€” Alex Ratner, Snorkel AI CEO (SiliconANGLE)

The money is earmarked for hiring engineers, funding AI-safety work, and building open-source benchmarks for evaluating models โ€” a tell that Snorkel wants to be seen as infrastructure for the whole field, not just a vendor to a handful of labs. The company says it expects to be profitable in 2026 (Reuters).

โš ๏ธ The Catch

An 18x year is spectacular and precarious in equal measure. A run rate is a snapshot โ€” a good month annualized โ€” not $375 million in the bank. And Snorkel's revenue is concentrated among a small set of very large customers: the same frontier labs that are all racing to build their own in-house data pipelines. If a couple of them decide to bring the work internal, the chart that justified $3.5 billion can bend the other way just as fast.

There's also the crowd. Mercor is bigger by revenue, Scale has Meta's balance sheet behind it, and Surge AI, Handshake, and Micro1 are all chasing the same expert-data dollars. Snorkel is betting its Stanford-bred methods and its focus on reinforcement-learning environments keep it a step ahead. That's a bet on execution, not a moat you can see from orbit.

๐ŸŽฏ What Happens Next

Watch two things. First, whether that run rate holds up as an actual annual number when the year closes โ€” the difference between "annualized" and "annual" is where a lot of AI valuations go to get a haircut. Second, the open-source benchmark work: if Snorkel becomes the group that defines how models get graded, it earns influence that's much harder to compete away than any single data contract.

Expect the rest of the sector to keep raising, too. When five companies in one niche are posting numbers between $375 million and $2 billion in run-rate revenue, the venture money follows, and fast.

๐Ÿงฉ Bigger Picture

For most of the last decade, the AI story was about the models. The quiet plot twist of 2026 is that the value is migrating to the supply chain โ€” the companies that feed the models. Meta spent $14.3 billion to lock up Scale. Snorkel went from a software license to a $3.5 billion data business in roughly a year. The labs still get the headlines, but the picks-and-shovels crowd is starting to get the margins.

Snorkel spent six years teaching machines to label data. Turns out the real business was selling the data itself โ€” and the industry can't buy it fast enough.


Sources