
Imagine a four-year-old startup that grows its revenue five times over in under eight months, without launching a flashy new product, just by feeding hungry AI models the human-verified data they need to get smarter. That’s exactly what happened to Micro1, whose gross annual run rate jumped from $100 million to $500 million between roughly December 2025 and August 2026, according to TechCrunch. The short answer: AI training data companies are booming because frontier AI labs have run out of easy internet text to train on and now need expert-verified, human-generated data, and startups like Micro1, Mercor, and Handshake are racing to supply it.
If you’re a student, fresher, or young professional in Odisha wondering where the real AI jobs are hiding beyond “prompt engineer,” this story is worth your attention. The people getting paid inside this boom aren’t just coders, they’re doctors, lawyers, scientists, and generalists teaching AI models how humans actually think and work.
What Is Micro1, and Why Is It Growing So Fast?
Micro1 is a data-labeling startup that pays domain experts, doctors, lawyers, scientists, and other specialists, on a contract basis to evaluate and improve AI model outputs. The company keeps roughly 60% to 70% of the revenue it bills to clients, which puts its net annual run rate somewhere between $150 million and $200 million even at a $500 million gross run rate.
Data labeling is the process of humans reviewing, correcting, ranking, or annotating information so an AI model can learn from it more accurately. Think of it as a teacher grading a student’s homework, except the “student” is a large language model, and the “homework” might be a legal contract, a medical diagnosis, or a snippet of code. Without this human feedback loop, AI models would keep repeating the same mistakes because raw internet text alone can’t teach nuance, accuracy, or professional judgment.
Did Micro1 always work in AI training data? No. Micro1 actually began life as an AI recruiting platform. Its founder, Ali Ansari, noticed that companies using Micro1’s tools to vet and hire engineers were actually clients looking for annotation talent, so he pivoted the business into data labeling as well, following a path similar to rival company Mercor.
The AI Training Data Boom, Why Is Demand Exploding Right Now?
The core reason AI training data companies are growing this fast is simple: AI labs have largely exhausted the “free” internet data available online, and the next leap in model performance now depends on carefully curated, expert-verified human data. Some researchers have even gone as far as predicting that future AI spending on data could eventually rival spending on computing infrastructure itself, a signal of just how central this resource has become.
Reinforcement learning gyms are structured environments where human experts evaluate an AI model’s responses and provide feedback that trains the model to improve, much like a training ground for an athlete. Micro1, for instance, uses domain experts to grade model outputs inside these “gyms,” building a feedback loop that gradually sharpens the AI’s reasoning and accuracy in specialized fields like law, medicine, and science.
Why does this matter for the wider AI industry? Because the quality of an AI model increasingly depends less on raw compute power and more on the quality of the data it learns from. This is why AI training data companies are being valued so aggressively by investors, they control a bottleneck resource that every major AI lab needs, whether that lab is building a chatbot, a coding assistant, or a robotics system.
Micro1 isn’t stopping at text and code either. The company is also building a robotics pre-training dataset by having hundreds of generalist workers record everyday object interactions inside their own homes, essentially teaching future robots what a doorknob, a spoon, or a laundry basket looks like from a human’s point of view.
What kind of work do these domain experts actually do day to day? A doctor might review an AI’s suggested diagnosis and flag where its reasoning breaks down. A lawyer might grade how well a model summarizes a contract clause or catches a legal risk. A software engineer might rate whether a model’s code actually compiles and solves the problem correctly. In every case, the expert isn’t writing content from scratch, they’re judging, correcting, and ranking AI-generated output so the model learns what “good” looks like in that specific profession.
This is very different from the more commonly discussed role of “prompt engineer,” which focuses on crafting inputs to get better outputs from an existing model. Data evaluation and reinforcement learning gym work instead focuses on judging and improving the model itself, which is why it tends to require deeper subject-matter expertise rather than just familiarity with AI tools.
Micro1 vs. Mercor vs. Handshake, Comparing the AI Data Labeling Giants
Micro1 is far from alone in this space. Here’s how it currently stacks up against its two best-known rivals among data labeling startups.
| Company | Gross Annualized Run Rate | Business Origin | Notable Detail |
| Mercor | ~$2 billion (as of mid-2026) | AI recruiting → data labeling | Largest of the three by revenue |
| Handshake | ~$1 billion (as of mid-2026) | Career/recruiting platform → data labeling | Reached $1B run rate earlier in 2026 |
| Micro1 | ~$500 million (as of August 2026) | AI recruiting → data labeling | Grew 5x in about eight months; raised Series A at a $500M valuation in September 2025 |
Even though Micro1 trails Mercor and Handshake in absolute revenue, its growth trajectory shows something more important for the industry as a whole: there’s enough demand for AI training data companies to support several large players simultaneously, not just a single winner-takes-all leader.
How Do AI Training Data Companies Actually Make Money?
Understanding the business model behind these data labeling startups helps explain why investors are so excited. Revenue generally comes from a mix of the following:
- Custom annotation contracts, clients pay for domain experts to label or evaluate specific datasets tailored to their model’s needs, and these contract sizes are reportedly growing at an accelerated pace for Micro1.
- Reinforcement learning gym services, labs pay for structured environments where experts grade and refine model outputs over time.
- Off-the-shelf synthetic data, pre-built datasets that can be sold to multiple customers at once, driving gross margins as high as 80% to 90% for this category, according to a person familiar with Micro1’s finances.
- Automated data generation, Micro1 is increasingly producing data without direct human involvement, such as generating automated descriptions of video content, which improves margins over time.
- Robotics and multimodal datasets, newer revenue lines built around real-world object interaction data for training robots and embodied AI systems.
This blend of human-expert labor and increasingly automated, reusable synthetic data generation is exactly why margins for AI training data companies are expected to expand rather than shrink as they scale.
Synthetic data generation is the process of creating new training examples using AI itself rather than relying purely on humans to produce them from scratch. Instead of paying a person to write out thousands of variations of a task, a company can generate many of them automatically and then use human experts only to review, correct, or approve the results. This hybrid approach, machines doing the heavy lifting of production while humans handle quality control, is a big part of why gross margins on “off-the-shelf” datasets can climb as high as 80% to 90%.
Why Are Investors So Interested in This Sector Right Now?
Venture investors are watching this space closely because these data-focused startups sit at a genuine bottleneck in the AI supply chain. Every major AI lab, whether it’s building a general-purpose chatbot, a specialized coding assistant, or a robotics model, ultimately needs a steady supply of high-quality, human-verified data to keep improving. That makes these startups less dependent on any single AI lab’s success and more like infrastructure providers to the entire industry.
The fact that three separate companies, Micro1, Mercor, and Handshake, can each report run rates in the hundreds of millions to billions of dollars within the same year suggests this isn’t a temporary spike tied to one product cycle. It looks more like a structural shift in how AI models get built, where human expertise is treated as a renewable resource that has to be continuously purchased, curated, and refined.
The China Data Controversy, Should Training Data Cross Borders?
Not everything about this boom is uncontroversial. Selling the same “off-the-shelf” datasets to multiple clients has triggered debate, with critics arguing that supplying this kind of data to Chinese AI developers could help their models close the performance gap with leading U.S. systems.
Does Micro1 sell its data to Chinese AI companies? According to founder Ali Ansari, no. Ansari stated on X last month that Micro1 does not sell data to Chinese model makers, adding that companies which do so are undermining claims of American AI dominance while “shameful[ly]” strengthening AI systems from countries considered adversarial competitors. He specifically pointed to rival model Kimi K3 as an example of the results of that practice.
This debate matters beyond U.S.–China competition, it’s a preview of the kind of data-sovereignty questions India will also need to navigate as AI training data companies scale locally and globally. Which experts get hired, which datasets get exported, and which countries benefit from that expert labor are all becoming live policy questions, not just business ones.
What This Boom Means for India and Odisha’s AI Talent Pipeline
Here’s the part that should matter most to students and young professionals reading this from Bhubaneswar or anywhere else in Odisha: this entire industry runs on domain experts, not just software engineers. Doctors, lawyers, scientists, linguists, and even generalists with strong reasoning skills are all being paid to train AI models through data labeling startups like Micro1, Mercor, and Handshake.
- If you have subject-matter expertise in any professional field, you are a candidate for reinforcement learning gym work, not just people with computer science degrees.
- Understanding how these firms operate, their workflows, quality checks, and evaluation frameworks, is becoming a genuinely marketable skill on its own.
- Freshers building career paths in applied AI should treat data evaluation, model grading, and annotation work as legitimate entry points into the AI economy, not lesser alternatives to “real” AI jobs.
As this sector scales globally, India’s large pool of English-speaking domain experts across medicine, law, engineering, and the sciences puts the country in a strong position to participate in this specific corner of the AI economy, provided the right training and awareness exists.
What Skills Should You Build to Get Into This Field?
Breaking into AI data evaluation work doesn’t require you to become a machine learning researcher overnight. Instead, focus on the intersection of your existing expertise and how AI models are actually judged and improved:
- Strengthen your core domain knowledge first. Whether that’s engineering, biology, law, finance, or another field, deep subject expertise is the actual product AI training data companies are paying for.
- Learn how large language models are evaluated. Understanding concepts like reinforcement learning gyms, hallucination detection, and output ranking will make you a stronger candidate for annotation and evaluation contracts.
- Practice structured, written feedback. Much of this work involves explaining precisely why an AI’s answer is right or wrong, not just marking it correct or incorrect, so clear written communication matters as much as technical knowledge.
- Stay current on which companies are hiring. Platforms tied to data labeling startups like Micro1, Mercor, and Handshake frequently open contract-based roles for specialists across many industries, not just computer science.
- Understand the ethics and geopolitics of data work. As the China data controversy shows, knowing where your labeled data ends up and who benefits from it is becoming part of doing this work responsibly.
None of this requires abandoning your primary career path. In fact, most of the experts fueling this boom are professionals with day jobs who take on contract-based annotation and evaluation work on the side.
FAQ: AI Training Data Companies and the Micro1 Growth Story
What is Micro1’s current valuation and revenue? Micro1 grew its gross annual run rate from $100 million to $500 million in about eight months as of August 2026, retaining roughly 60% to 70% of that as net revenue. The company raised its Series A at a $500 million valuation in September 2025 and, per TechCrunch, may have recently raised another round at a significantly higher valuation.
How does Micro1 compare to Mercor and Handshake? Micro1 is smaller than both rivals, Mercor reached roughly $2 billion in gross annualized revenue and Handshake reached about $1 billion in 2026, but Micro1’s rapid five-fold growth shows the overall market for this sector is large enough to support multiple major players.
What exactly do data labeling startups do? They hire contract-based domain experts, such as doctors, lawyers, and scientists, to evaluate, correct, and rank AI model outputs. This human feedback trains models to be more accurate, especially in specialized professional fields where mistakes carry real consequences.
What is a reinforcement learning gym in simple terms? It’s a structured setup where human experts repeatedly evaluate and grade an AI model’s answers, and that feedback is used to fine-tune the model’s future responses, similar to a coach reviewing an athlete’s performance after every match.
Why are AI training data companies growing faster than expected? Because internet-scale training data is largely exhausted, and further AI improvement now depends heavily on expert-verified human data, synthetic data generation, and reinforcement learning feedback loops rather than just adding more raw compute.
Is training data for AI models a legitimate career path in India? Yes. Domain experts and generalists alike are being paid by AI training data companies to evaluate, annotate, and improve AI systems, and this demand is expected to keep growing as more labs compete for high-quality human feedback.
Do I need a computer science background to work in AI data evaluation? No. Most of the highest-paid roles inside AI training data companies go to doctors, lawyers, scientists, and other domain specialists, not software engineers, because the work is about judging accuracy and reasoning within a profession rather than writing code.
What’s the difference between data labeling and synthetic data generation? Data labeling relies on humans reviewing or annotating existing information, while synthetic data generation uses AI to automatically create new training examples, which humans then review or correct. Companies like Micro1 increasingly blend both approaches to improve speed and margins.
Curious how you could actually break into AI data evaluation, annotation, or reinforcement learning gym work as a student or fresher in Odisha? Explore Kalinga.ai’s AI training programs and job listings to see how this global boom connects to real opportunities closer to home.