Micro1, a four-year-old startup supplying training data to AI labs and corporations, has grown its gross annual run rate from $100 million to $500 million over the past eight months, according to a person familiar with the company. The startup retains roughly 60% to 70% of that figure, putting its net annual run rate between $150 million and $200 million.
The surge reflects a broader boom in the AI data-labeling sector, where demand for unique, high-quality training data has become nearly bottomless. Micro1’s growth, while significant, still trails larger competitors like Mercor, which hit $2 billion in gross annualized revenue this summer, and Handshake, which reached $1 billion earlier this year. Yet the fact that multiple players are scaling rapidly suggests the market can support several major suppliers.
Also read: Can Urine Cool Data Centers? The Real Science Behind Liquid Death's Viral Ad
From recruiting to data labeling
Micro1 began as an AI recruiting platform, similar to Mercor. Founder Ali Ansari pivoted to data labeling after noticing that clients were using his platform to vet and recruit engineers for annotation work. The company now employs domain experts—doctors, lawyers, scientists—on a contract basis to evaluate model outputs, a process known as reinforcement learning gyms.
Beyond human-in-the-loop services, Micro1 is increasingly generating synthetic data without human involvement, such as automated descriptions of video content. Some of this off-the-shelf data can be sold to multiple customers, driving gross margins as high as 80% to 90%, according to a person familiar with the startup’s finances.
Also read: Ramp launches Router, an AI model routing service for enterprises
Synthetic data and geopolitical concerns
The practice of selling the same datasets to multiple clients has sparked controversy. Critics argue that distributing off-the-shelf data to Chinese AI developers helps their models match the capabilities of top U.S. models. Ansari addressed this directly on X last month, stating that Micro1 does not sell data to Chinese model makers.
“Some human data companies work with foreign adversaries. and the results show today in Kimi K3,” Ansari posted. “We believe it’s shameful to claim American AI dominance desires while selling millions worth of data to countries that we are in adversarial competition with.”
What’s next for Micro1
The startup is also building a robotics pre-training dataset, with hundreds of generalists recording everyday object interactions in their homes. Contract sizes are growing at an accelerating pace, and management expects margins to expand over time as synthetic data becomes a larger share of revenue.
Micro1 raised its Series A at a $500 million valuation last September. TechCrunch understands the startup may have recently raised another round at a significantly higher valuation, though Micro1 did not respond to a request for comment.
Some researchers hypothesize that future AI spending on data could rival spending on compute, which would bode well for Micro1 and its peers. The startup’s trajectory illustrates how the AI boom is creating a lucrative niche for companies that can source, curate, and package the raw material that powers model training.
This article is for informational purposes only and does not constitute financial advice. The AI data market is volatile and subject to rapid change; readers should conduct their own research before making any investment decisions.

Be the first to comment