Market entry assessment, September 2026: six candidate categories, one recommendation.
By Mohinish Shaikh and Pragadeesh VS
Published: September 7, 2026 · Read time: ~14 minutes
The commodity labeling tier is being automated toward zero. The layer above it, reinforcement-learning environments with working verifiers, is the only high-value category where the deliverable is software, the buyer pays per artifact rather than per hour, and two engineers in India can compete on craft instead of headcount.
Recommended entry: RL environments with verifiers. Scope: which of six high-value data categories a two-person, engineering-led startup with an India-based workforce should enter. All figures dated and sourced. Several are estimates or single-source; see the reliability notes at the end.
Demand is enormous and highly concentrated, and it has split cleanly in two: the bottom is collapsing in price while the top is starved of supply.
Annual human-data spend per frontier lab (Time, 2025 investigation)
Of category revenue held by Scale, Surge, Mercor and Handshake (Deedy Das venture map, via Pebblous)
Of routine labels now correctly pre-filled by foundation models, compressing that tier (Shaip)
Cost advantage of India delivery over US teams for equivalent work (Precise BPO)
Synthetic data is a complement rather than a replacement: cited optimal mixes sit around 60–70% human to 30–40% synthetic, and pure-synthetic training risks model collapse. Machines now handle volume; humans set policy, audit edges, and build the environments the machines train inside. That is where the money moved.
Six categories measured on the axes that decide whether two engineers can actually win: how badly buyers want it, whether a small team can hold a position, how much capital it eats, how fast it pays, and whether it can be delivered from India without a jurisdictional blocker.
Software deliverable, per-artifact pricing, non-personal data, engineering is the moat. Enter here.
Natural bolt-on to environments. Build synthetic capture; keep real-user personal data out of scope.
Biggest revenue pool, but the moat is a 30,000-expert network you cannot rebuild. Viable only as a narrow vertical.
Automating fast, and the premium tier is gated to US persons. Skip as a primary business.
Home-turf advantage, but the buyer is a thin domestic budget reached through slow government tenders.
Real demand, wrong profile: capital-heavy, ops-heavy, geographically anchored, slow. Avoid.
Each gauge is five segments; more filled is better. Scores are judgements drawn from the evidence in the dossiers below, not published metrics.
What actually ships, who buys it, what it costs, who already owns the position, and what would stop you delivering it from India.
An action space plus surrounding state (file systems, simulated apps, environment variables), usually delivered as a Docker container, plus tasks (a prompt and a grader). Graders are unit tests, state inspection, programmatic checks, or LLM-judges against rubrics. A commercial example: a repo snapshot, task spec, and a verifier scoring each attempt on a continuous 0–1 reward, shipped as "the container and the reward entrypoint," plus reward-labeled agent trajectories. Coding environments usually bundle one task each; a computer-use replica (an Excel, Bloomberg or Airbnb clone) can carry hundreds.
Frontier labs (Anthropic; OpenAI and xAI named via job postings), neolabs such as Cursor, and product partnerships (Benchling with Anthropic for biology; OpenAI with Shopify and Stripe for shopping). Environment vendors also subcontract build capacity overseas.
Contracts often six to seven figures per quarter; deal sizes of $300k–$500k cited by one neolab; website replicas around $20k; a complex product replica such as Slack around $300k; per-task $200–$2,000, rarely to $20k for hard software engineering. Exclusive deals cost 4–5× non-exclusive. (Epoch AI survey of 18 insiders, Jan 2026; SemiAnalysis)
The Information reported in September 2025 that Anthropic discussed spending "at the level of $1 billion per year on training environments alone." For scale, OpenAI's R&D consumed roughly $19B of about $34B total 2025 spending, and Greg Brockman testified it would spend around $50B on compute in 2026, triple 2025.
Specialists: Mechanize (SF, ~$9.1M raised, working with Anthropic, $500k engineer salaries), Halluminate (YC S25, narrowed to financial-services environments, ~$1.3M ARR reported 2025), Prime Intellect (open Environments Hub, 500+ community environments, Verifiers library), HUD, and others catalogued on rl-list.com. The large data vendors (Mercor, Surge, Turing, Handshake) are moving in with distribution, but it is an open question whether labeling operations translate into sound verifier engineering. The field is fragmented and early.
Capital: negligible. Skill: high. Reward-hacking resistance takes many iterations; difficulty must be calibrated to roughly a 2–3% minimum pass rate with a smooth gradient and compositional skills. Every interviewee named the same hard problem: scaling task volume without quality collapse. That is a management problem a software-minded team can partly convert into a tooling problem, which is exactly the opening.
Low, and inverted: environments are themselves the machinery that generates synthetic data, so the synthetic shift increases demand here. Real risks are labs in-housing (already "substantially more in-housing") and a flood of vibe-coded clones at the low end: "a large amount of useless bad environments out there."
Clear. Environments are non-personal software artifacts, outside the DPDP Act's personal-data scope. Commercial frontier work is open to India-based delivery, and environment vendors already hire overseas developers to replicate site UIs. A US wrapper additionally opens anything government-adjacent.
A natural bolt-on to environments rather than a separate business: the same engineering that builds an environment also produces the recorded trajectories labs want. Build synthetic capture rather than instrumenting real users, and keep real-user personal data out of scope entirely, because that is what keeps the category clean under the DPDP Act and simple to sell across borders. Full dossier in the interactive assessment.
Credentialed experts author original problems, worked solutions, grading rubrics, evaluations and reasoning traces: production, not labeling. Domains in demand: medicine, law, senior software engineering, advanced mathematics, quantitative finance, accounting, PhD sciences.
Mercor, Surge, Handshake, Micro1, Turing and AfterQuery place vetted experts onto lab projects. Mercor runs about a 35% take rate; The Information, citing internal documents, put gross revenue at roughly $614M in H1 2026 and a $2B annualized gross run-rate by June (up from about $760M at end-2025 and $500M in September 2025), with 91% of revenue from foundation-model companies such as OpenAI and Anthropic. Sacra estimates contractors keep 60–70% of top line, implying H1 net revenue of roughly $180M–$250M. It distributes about $1.5M per day to 30,000+ contractors across 45+ countries, India its largest talent source.
Mercor average around $85/hr (advertised $81–$114). Tiers: $12–$25/hr entry generalist; $25–$53/hr specialized RLHF, coding and translation; $75–$200+/hr credentialed experts. Medical MDs $100–$210/hr with a top decile near $250/hr; offensive-security $200–$250/hr; law, finance and quant in between. (Contractor-reported, 2026)
Sacra estimates Handshake reached $1.1B annualized gross revenue in April 2026, up 349% year over year from about $245M, with roughly $450M net after contractor payouts; The Information put its AI-training gross ARR near $1B, up from $550M in January 2026.
The moat is the credentialed-expert network and the vetting engine (Mercor's APEX AI interview), not technology, and not replicable by two engineers. You would be reselling the same India expert labour the incumbents already recruit directly, with no differentiation.
Open but undifferentiated. Only worth pursuing as a jurisdiction-specific vertical where you hold real advantage (Indian law, ICAI accounting, Indian medical boards) rather than head-on.
Automating fast, and the premium tier (the defence- and government-adjacent work that pays $200–$250 an expert hour) is gated to US persons. That combination closes the high end to an India-based team while the low end erodes. Skip it as a primary business; revisit only as an add-on to environment work. Full dossier in the interactive assessment.
Genuine home-turf advantage, but the buyer is a thin domestic budget reached through slow government tenders (IndiaAI Mission, CPP portal empanelment, BharatGen, Sarvam). Defensibility is real; the ceiling and the sales cycle are the problem. A reasonable second line once delivery credibility exists, not a first revenue source. Full dossier in the interactive assessment.
Real demand, wrong profile. Teleoperation episodes price at $8–$40 each, and earning that requires rigs, physical space, operators, and geographic anchoring: capital-heavy, ops-heavy and slow, the exact inverse of a two-person software team's advantages. Avoid. Full dossier in the interactive assessment.
The single most important structural fact for an India-based team: categories priced per artifact preserve the cost arbitrage, categories priced per expert hour have already competed it away.
| Unit | Price | Category | Source |
|---|---|---|---|
| Complex product replica | ~$300,000 | RL environment (e.g. Slack clone) | Epoch AI, Jan 2026 |
| Quarterly contract | $300k–$500k | RL environments, neolab deal size | Epoch AI, Jan 2026 |
| Website replica ("UI gym") | ~$20,000 | RL environment | Epoch AI / SemiAnalysis |
| Single task | $200–$2,000 | RL environment task (rarely to $20k) | Epoch AI, Jan 2026 |
| Exclusivity premium | 4–5× | RL environments | Epoch AI, Jan 2026 |
| Compute burned per task | ~$2,400 | RL training, the reason labs pay up | Mechanize estimate |
| Expert hour, medical MD | $100–$250 | Expert data generation | Contractor-reported, 2026 |
| Expert hour, offensive security | $200–$250 | Expert data / red-teaming | Contractor-reported, 2026 |
| Expert hour, marketplace average | ~$85 | Mercor blended rate | Contractor-reported, 2026 |
| Red-team finding | up to $100,000 | Bug bounty maximum | OpenAI programme |
| Teleoperation episode | $8–$40 | Robotics data | Robotics Center of Silicon Valley; Dexset |
| India commodity labeling hour | $5–$10 | The tier to stay out of | Market rate, 2026 |
| India salaried annotator, annual | ~₹9.3 lakh | ≈ $11,000 | SalaryExpert, 2026 |
The one hard net-margin figure available in this market is sobering, and it should shape the whole business model.
Invisible Technologies EBITDA
Mercor take rate on gross
Contractor share of marketplace top line
India cost advantage vs US delivery
Invisible Technologies runs roughly 11% EBITDA on about $134M of revenue (Sacra). Human-in-the-loop is a thin-margin business unless automation is layered on top. Note too that headline "ARR" figures across this sector are usually gross marketplace volume before contractor payouts. Turing's economics are explicitly described as a staffing spread, earning the difference between what clients pay and what talent receives.
Your arbitrage survives only where pricing is per artifact. An India engineering team building environments sold at $20k apiece into US per-environment pricing captures the gap. The same team reselling expert hours does not.
The clearest structural opening in the market: environment companies already hire overseas developers to replicate site UIs (SemiAnalysis), and OpenAI reportedly bought hundreds of roughly $20k UI gyms. Be that development shop: for environment vendors first, then directly for labs.
Why: Fastest to revenue, engineering-native, non-personal data, no expert network required. Price per environment, not per hour.
Agent startups and neolabs need training and evaluation environments but cannot build their own. Smaller deals, but you own the customer relationship and the roadmap.
Why: Fallback if labs accelerate in-housing, and the natural home for a vertical specialisation later.
IndiaAI Mission RFEs, CPP portal empanelment, and direct contracts with BharatGen, Sarvam and others.
Why not first: Slow tender cycles and thin budgets make this a poor first revenue source, though a reasonable second line once you have delivery credibility.
Structure the company as a US wrapper over India engineering, the Turing pattern (Palo Alto incorporation, India-heavy delivery). India-domiciled entities do not appear to win direct frontier RL contracts at scale today, and the wrapper also opens government-adjacent work later. iMerit made the same move: India-founded, now headquartered in Los Gatos, climbing from annotation into frontier expert-data tiers.
Gate to advance: two environments adopted or downloaded, or one paid bounty, within eight weeks.
Gate to advance: first paid environment contract of $20k–$100k by month six. No contract by then means switch to path 2 before considering a category change.
Gate to advance: a recurring or exclusive environment contract. Exclusivity prices at 4–5× non-exclusive.
Compiled September 2026 from Epoch AI, SemiAnalysis, Sacra, The Information, Time, Contrary Research, Pebblous, The Business Research Company, Counterpoint, SalaryExpert, Precise BPO, Shaip, AI4Bharat, and company and government sources including MeitY and the IndiaAI Mission.
Mohinish Shaikh and Pragadeesh VS work on AI software, media and investing at Serverlessvc.com.
Prepared as a market-entry assessment. Not investment advice. Figures are dated and, in several cases, estimates or single-source; see the reliability notes above.