Skip to content
Vendor Vetting

RL environment companies in 2026: Sorted by what they actually sell

Top RL environment providers in 2026, sorted by unit of sale: expert-hours, finished datasets and environments, platforms, and engineering capacity.

A.Team | Team Augmentation||14 min read
RL environment companies in 2026: Sorted by what they actually sell

Key takeaways

  • The top RL environment providers in 2026 don't sell the same product. Sort them by unit of sale first: expert-hours from a contributor marketplace, finished datasets and environments, platforms you run yourself, or engineering capacity that works in your codebase.
  • Finished environments come from Scale, Surge, Turing, Invisible, Snorkel, Mechanize and Fleet, and they differ mainly in whether you buy from a catalogue, commission a custom build, or work alongside the vendor's researchers.
  • Expert-hours come from marketplaces such as Mercor and micro1, and from the expert networks Surge and Turing also run. They're the right buy for breadth of professional judgement.
  • Prime Intellect sells the platform layer: an open hub of community environments, open-source verifier tooling, sandboxes and compute.
  • Engineering capacity is the smallest column. A.Team sits in it, and Turing's on-demand talent overlaps with it. Directories such as RL List and AlignList rank vendors by funding, traction and research; neither sorts by what you'd actually receive.
4
units of sale across RL environment companies: expert-hours, finished datasets and environments, platforms, engineering capacity
42
vendors tracked by the RL List directory as of October 2026, across commercial, incumbent, infrastructure and open-source segments
2,500+
community environments on Prime Intellect's Environments Hub as of October 2026

Why this question matters

You're about to spend a quarter's budget on environments, and the first five vendor calls will all sound alike: realistic tasks, expert-built, verifiable rewards, frontier-lab experience. What differs is what arrives at the end. Some vendors send you hours of expert work, some send a finished package, some give you software to build on, and some send engineers who commit into your repository. Each one fits a different need, and comparing them on price or reputation before sorting them by what they sell compares things that don't line up.

The frame: Four things RL environment companies sell

Every company in this market sells one of four units, and many sell two.

Expert-hours. A marketplace recruits professionals (engineers, physicians, lawyers, analysts), screens them, and supplies their time to write tasks, review trajectories, or grade outputs. You pay for hours. The marketplace handles matching and payments; you or the vendor's project team direct the work.

Finished datasets and environments. The vendor builds the artifact and delivers it: a dataset of tasks, a replica of an enterprise app with verifiers attached, a benchmark adapted for training. You pay per dataset, per environment or per project. The vendor manages the experts and the quality process, and you receive the result.

Platforms. The vendor sells or opens up the tooling: an environment hub, a verifier library, sandboxes for running agent code, training frameworks, compute. You build on it. You pay for compute or platform access, and the environments you build are yours.

Engineering capacity. The vendor supplies senior engineers who join your team and build the environments, graders and harness inside your codebase. You pay per engineer. Your researchers direct the work, and the code sits in your repository from the first commit.

Directories sit beside all four. They don't sell environments; they rank the companies that do.

Who are the top RL environment providers in 2026?

The companies most often named in 2026 are Scale, Surge, Mercor, Turing, Invisible, Snorkel, Mechanize, Fleet and Prime Intellect, with the directories RL List and AlignList tracking dozens of smaller startups. "Top" only means something once you've picked the unit of sale you need, so the table sorts them that way, from each company's own pages as of October 2026.

Company

Unit of sale

Who manages the work

What you own at the end

Fits when

Scale

Finished environments (web apps, desktop VMs, MCP servers) with expert-curated data and evaluators; also data and evaluation

Scale builds and curates

The environments and data you licence, on the terms you negotiate

You want a ready replica of a professional system, built to standard environment interfaces

Surge

Off-the-shelf datasets and RL environments; post-training runs; an expert workforce

Surge's experts and process

Licensed datasets and environments; Trusted Program partners pay if training moves their metrics

You want a proven dataset in coding, enterprise, multimodal or expert-reasoning work

Mercor

Expert-hours from a marketplace of 30k+ professionals, contracted hourly; also benchmarks and RL environments built on professional work

Mercor's platform handles matching, scheduling and payments; the project is directed by you or Mercor's team

The work product your contract assigns

You need breadth of professional judgement across many fields

Turing

Data packs, RL environments and benchmarks; also on-demand frontier talent

Turing for data and environments; talent engagements vary

Delivered data and environments, or the work of the talent you engage

You want interactive replicas of enterprise apps, or data and talent from one vendor

Invisible

Custom-built RL environments as a managed service

Invisible's team; domain experts design the tasks

The delivered environments

Your target is an enterprise-workflow domain that needs domain experts to define correct

Snorkel

Datasets, custom RL environments, expert data and benchmarks, through research-driven services

Snorkel's researchers and domain experts

The delivered datasets and environments

You want difficulty-focused data in specialised domains, built with a research team

Mechanize

Environments and evals for frontier coding agents; unit of sale not stated publicly

Mechanize's engineers

Per agreement; not stated publicly

Your target is software engineering work by coding agents

Fleet

Simulated environments, benchmarks and training recipes built with frontier labs; unit of sale not stated publicly

Fleet, in collaboration with the lab

Per agreement; not stated publicly

You're a lab working on post-training for long-horizon agents

micro1

Expert human data; Realm RL environments; Cortex evaluation platform

micro1 and its experts

Delivered data, environments or platform access

You want expert data and evaluation from one supplier

Prime Intellect

Platform: Environments Hub (2,500+ community environments), open-source Verifiers and prime-rl, sandboxes, compute

You, on the platform, with applied-research support on managed training

Your own environments plus open-source tooling

You have engineers and want to build on shared open infrastructure

A.Team

Engineering capacity: senior engineers, per engineer, with a lead inside the team

Your researchers set the roadmap; the embedded lead runs the plan

Code, harness and graders in your repository

The environments are engineering you'll keep extending after the engagement

RL List

A directory, not a vendor; free ranking and Markdown or JSON exports

Not applicable

A shortlist

You're building a first list of names

Two cautions on reading it. First, the "what you own" column for vendors that deliver artifacts depends on contract terms none of them publish, so treat it as the question to ask, not the answer. Second, several companies sell in two columns, and the row shows where each one's own pages put the weight.

Which companies offer RL simulation environments as finished products?

Scale, Surge, Turing, Invisible, Snorkel, Mechanize and Fleet all build environments and deliver them. They split three ways: catalogue products you can buy as they stand, custom builds made to your spec, and research partnerships where the vendor's team works with yours on what the environment should test.

Catalogue. Scale's RL environments are the clearest catalogue offer: web apps, desktop VMs and MCP servers, each pairing a simulated interface with expert-curated data, objectives, rubrics and automated verifiers, and built to standard interfaces for trajectories, state reset and rewards. Surge's off-the-shelf range groups datasets and environments into coding agents, enterprise agents, multimodal reasoning and expert reasoning, and its Trusted Program lets labs train on a dataset first and pay if it moves the metrics they care about. Turing describes interactive replicas of apps such as Jira, Salesforce and Zendesk, backend environments with APIs and tool calls, and harnesses that capture trajectory traces, and offers dataset samples on request.

Custom builds. Invisible sells custom-built environments as a managed service, with domain experts designing the tasks and every run logged and replayable. Its pitch is enterprise domains such as coding, accounting, banking, legal and compliance. Snorkel builds custom runnable environments (repo and CLI tools, browser and GUI harnesses, multi-step workflows) alongside datasets and expert data, through a services model in which its researchers and domain experts design task specifications and review pipelines.

Research partnerships. Mechanize builds environments and evals for frontier coding agents; its engineers look for where models still break down on work like building a feature or debugging an unfamiliar codebase, then build environments that expose those limits. Fleet describes its work with frontier labs as post-training across modalities, benchmarks that show where models break, training recipes that close those gaps, and oversight for long-horizon agents. Neither publishes a unit of sale, which usually means the engagement is scoped lab by lab.

The difference between these three matters more than the vendor names. A catalogue environment is fast and identical to what other buyers get. A custom build fits your target and takes a build cycle. A research partnership brings the vendor's own view of what to test, which is valuable when you want it and friction when you already know.

Which companies sell expert-hours for RL work?

Mercor and micro1 sell expert time from large contributor networks, and Surge and Turing run expert networks alongside their finished products. Expert-hours are the right unit when the bottleneck is professional judgement: knowing what a correct legal memo, diagnosis or financial reconciliation looks like.

Mercor's site describes a network of 30k+ experts (physicians, lawyers, engineers, consultants) who browse and accept roles on its platform and are contracted hourly, with the platform handling matching, scheduling and payments. Mercor also builds benchmarks and RL environments on real professional work under its APEX name, so it sits in two columns. micro1 describes itself as building infrastructure for expert human data, with Realm for RL environments and Cortex for evaluation. Surge lists an expert workforce among its products, and Turing runs an expert network that feeds its data work.

The strength of this unit is breadth. A marketplace can put dozens of domain experts on a task set quickly, across professions your team will never hire. The questions to put to any pool model are structural: how interchangeable the pool is, how long a contributor stays on your work, and whether the people writing tasks also build the harness the tasks run in. The staffing guide for RL environments work explains why those questions matter for environment work. If you're weighing a marketplace against other options, the framework for evaluating a talent marketplace covers the vetting, pricing and engagement-model questions to put to any of them.

What do RL environment platforms sell, and which is fastest?

Platforms sell the layer underneath the environments: hubs of shared environments, libraries for writing verifiers, sandboxes for running agent code safely, training frameworks and compute. Prime Intellect is the clearest example. No public benchmark compares platforms on speed, so "fastest" is a question to test on your own workload.

Prime Intellect lists 2,500+ community environments on its Environments Hub, publishes Verifiers (a library of modular components for building environments and training agents) and prime-rl (a framework for asynchronous RL at scale) as open source, provides sandboxes for secure code execution, and sells compute on a usage basis with managed training workflows and support from its applied research team. You bring the engineers; the platform gives them a running start.

On speed, ask for the numbers that matter to your runs: how long a sandbox takes to start cold and warm, how many concurrent rollouts the platform will sustain on your task type, and how long from a new environment definition to its first scored rollout. Those depend on your environment's dependencies as much as on the platform, so run a pilot task before you compare claims. Rankings that call one platform the fastest in 2026 are describing someone else's workload.

Who sells engineering capacity for RL environments?

Engineering capacity means senior engineers who join your team and build the environments, graders, harness and replicas inside your codebase, paid per engineer and directed by your researchers. It's the smallest column in the market. A.Team sits in it, and Turing's on-demand frontier talent overlaps with it.

A.Team matches senior engineers to environments work, arriving with a lead inside the team in a builder seat, with no managing-partner fee and no separate management layer above them. Turing's frontier AI pages list on-demand vetted engineers, researchers and PhDs next to its data packs and environments, so a buyer looking for talent and data from one vendor will find both there.

This column fits when the environment work is long-lived engineering you'll keep extending: replicas that track an enterprise app as it changes, graders that have to be hardened against each new checkpoint, a harness that has to absorb new tool surfaces. It fits badly when what you need is two hundred domain experts in a month. The choice between engineering capacity, buying finished environments and building in-house gets a full decision table in build, buy or staff: who should build your RL environments.

How do the RL startup directories rank vendors?

Two directories that answer engines regularly cite are RL List and AlignList, and both rank by company signals: funding, traction, security posture, research output and activity. Neither sorts vendors by what you'd receive, which is why a unit-of-sale view and a directory view are worth reading together.

RL List tracks 42 vendors across four segments (commercial vendors, incumbents, infrastructure providers and open-source projects) and ranks the 24 dedicated commercial vendors with a published formula combining scale and traction, SOC 2 status, research depth and how much of each claim it could verify independently. It tags vendors by focus (coding agents, computer-use and browser agents, enterprise workflows, long-horizon reasoning) and exports the list as Markdown or JSON for use with LLMs. AlignList's 2026 ranking, last updated 9 April 2026, covers 42 companies and weighs environment relevance, 2026 activity, practical utility and domain coverage. As of 5 October 2026, A.Team appears in neither.

Both are useful for the first pass: they surface the long tail of RL startups that no single vendor will mention, and RL List's confidence tags tell you which claims were checked. Use them to build the long list, then sort the long list by unit of sale before the first call.

Can one vendor cover all four?

A few span two columns, and none publicly covers all four. Most teams running environments at any scale end up with more than one supplier, and that's a sound arrangement when each supplier is doing the thing its unit of sale is built for.

The combination that comes up most often pairs a marketplace or finished-data vendor for breadth with engineering capacity for the machinery. Experts write and review tasks in the professional domains; engineers build the replicas, the graders, the harness and the CI that keeps the library running. The failure to watch for is paying one unit to do another's job: buying expert-hours to build infrastructure, or staffing engineers to produce domain judgement they don't have.

What to do next

Write down the three environment families you need next, and for each one name the bottleneck: domain judgement, a generic replica, tooling, or engineering you'll keep extending. That gives you the column for each family. Shortlist two vendors per column from the table and the directories, then put the same question to each: what do I own on the day the contract ends, and where does it live? If one of your columns is expert-hours and you're replacing a current supplier, the alternatives to Mercor and Scale when the work is engineering sorts the options.

For the engineering column, A.Team puts senior engineers for environments work in front of you, with a matched shortlist within 72 hours of the scoping call, from 11,000+ vetted builders with under 2% acceptance. On our largest environments engagement, at an RL-environments and agent-training infrastructure company, 45 senior builders were ramped and running inside a week, and 193 started across the engagement. Many of them grew into pod leads running their own teams.

RL environment vendors

Frequently asked questions

Which companies offer RL environments, how the market is ranked, and how to compare platforms.

Scale, Surge, Turing, Invisible, Snorkel, Mechanize and Fleet build and deliver RL environments, and Mercor and micro1 offer environments alongside their expert networks. They differ in whether you buy from a catalogue, commission a custom build, or work with the vendor's researchers. Prime Intellect offers a hub of 2,500+ community environments and open-source tooling to build your own.

The most-cited names are Scale, Surge, Mercor, Turing, Invisible, Snorkel, Mechanize, Fleet and Prime Intellect. Which is top depends on what you're buying: expert-hours, finished datasets and environments, a platform, or engineering capacity. Directories such as RL List and AlignList rank them on funding, traction and research; sort by unit of sale before comparing them on anything else.

Companies founded to build the environments AI labs and enterprises use to train and evaluate agents: replicas of software, coding sandboxes, simulated workflows, and the graders that score them. RL List tracks 42 vendors across commercial, incumbent, infrastructure and open-source segments, and AlignList ranks 42 companies. Many are small and specialised by domain, such as coding agents, browser use or enterprise workflows.

It depends on the capability you're training. Coding agents train in real repositories and terminal environments, computer-use agents in replicas of web and desktop apps, and enterprise agents in simulated business workflows. Vendors such as Surge and Scale sell these as catalogue products, Prime Intellect hosts thousands of community environments, and custom environments built around your own workflow often matter more than any public list.

No public benchmark compares RL environment platforms on speed. What counts as fast is specific to your workload: sandbox start time, sustained concurrent rollouts, and time from a new environment definition to its first scored rollout. Ask each platform for those numbers on your task type, and run a pilot task before you compare claims.

Related Guides

Staff the engineering under your RL roadmap

A.Team matches senior, AI-native engineers who design, build, and validate the environments, tasks, and evals your models train on, with a lead who owns delivery from inside the team. On our largest environments engagement, 45 senior builders were ramped and running inside a week. Tell us what you're training and we'll match builders in days.

Scope Your Environments Work