Build, buy or staff: Who should build your RL environments
Where can I get custom RL environments for enterprise workflows? Build in-house, buy finished tasks, or staff senior engineers under your roadmap.

Key takeaways
- Where can I get custom RL environments for enterprise workflows? From three places: your own team, a vendor that sells finished tasks or environments, or senior engineers staffed under your own roadmap with a lead embedded among them. Most published guidance compares only the first two.
- Build in-house when the environment is part of the research itself, and when your engineers already know the domain the way a practitioner does. Coding and maths usually qualify.
- Buy finished tasks or environments when you need breadth across domains your team can't cover, or a first dataset fast. You get an artifact on the vendor's schedule, under the vendor's process.
- Staff senior engineers when the work is the machinery around the tasks (harness, graders, provisioning, replicas of enterprise apps) and you need to own it and keep extending it after the engagement ends.
- The question that sorts the three is what happens when the model gets better. Tasks saturate, graders get gamed in new ways, and somebody has to rebuild. Decide now who that somebody is.
Why this question matters
You've got a model that needs to get better at work that happens inside business software: tickets moving through a queue, records updated in a CRM, a reconciliation run across three systems. That takes environments built around those workflows, and the build is bigger than it looks from the roadmap. The route you pick decides who holds the environment library in a year, how fast you can change it, and what you're left with when the vendor contract or the hiring plan runs out.
The frame: Three routes, three different things you're paying for
Each route sells a different unit, and the unit predicts how the engagement behaves better than anything in the pitch.
Build in-house. You hire engineers onto your payroll and they build the environments, the harness, and the graders inside your codebase. You're paying for headcount. You get full control and full ownership, and you carry the hiring time, the management load, and the fixed cost through the quiet months between training runs.
Buy finished tasks or environments. A vendor delivers the artifact: a dataset of tasks, a set of environments, a benchmark, sometimes a hosted replica of an enterprise app with verifiers attached. You're paying per task, per environment, or per dataset. The vendor recruits the experts, runs the quality process, and hands you the result.
Staff senior engineers under your roadmap. A team of senior engineers joins your work, commits into your repository, and takes direction from your researchers, with a senior lead inside the team who owns the plan and the quality bar from a builder seat. You're paying per engineer. You get capacity that behaves like your own team without the hiring cycle, and the code stays with you.
The incumbents mostly write about the first two. Epoch's FAQ on RL environments, built on interviews with 18 people across startups and labs, sorts supply into specialised environment startups, data vendors expanding into environments, in-house lab teams, and product companies partnering with labs. Invisible's enterprise guide frames the decision as building internally against procuring externally. Neither treats staffed engineering capacity as its own route, and it's the one that fits a lot of enterprise-workflow work best.
Where can I get custom RL environments for enterprise workflows?
Finished enterprise environments come from vendors such as Scale and Turing, custom environment builds come from vendors such as Invisible, and the engineering to build and run your own comes from your team or a staffed one. Which one fits depends on how specific the workflow is to you, and how often it'll change.
The off-the-shelf end is real and worth checking first. Scale's RL environments come as web apps, desktop VMs, and MCP servers, each pairing a simulated interface with expert-curated data and evaluators, built to standard environment interfaces. Turing describes interactive replicas of apps such as Jira, Salesforce, and Zendesk, backend environments with APIs and tool calls, and harnesses that capture trajectory traces. Surge's off-the-shelf catalogue includes an enterprise-agents category of multi-step business workflows. For a widely used SaaS tool and a generic workflow, one of these may reach a first training run faster than anything you'd build.
"Custom" is where it changes. Invisible makes the point itself: an agent that handles your workflows, your data formats, and your quality standards needs an environment built around your operations. That means a replica seeded with data that looks like yours, tools that behave like your internal APIs, and a grader that checks the end state your operators would actually accept. Somebody has to learn your workflow well enough to encode it, and that somebody has to still be around when the workflow changes or the model starts passing everything.
So the real choice for custom enterprise work is who does the learning. A vendor building a custom environment learns your workflow inside its own team and hands back the result. A staffed team learns it inside yours. Your own hires learn it slowest and keep it longest.
Build, buy or staff: How do the three compare?
Read the table by the row you can least afford to get wrong. For most teams that's the last two.
Dimension | Build in-house | Buy finished tasks or environments | Staff senior engineers under your roadmap |
|---|---|---|---|
What you pay for | Salaried headcount | A task, an environment, or a dataset | An engineer's time, per engineer |
Time to ramp | The length of a senior hiring cycle per seat | Fast for catalogue items; a scoping and build cycle for custom work | Days to a shortlist, then the team's ramp on your codebase |
Who directs the work | Your researchers and engineering leads | The vendor's process, against your spec | Your researchers set the roadmap; a lead inside the team runs the plan |
Who manages the people | You | The vendor | The embedded lead day to day, with no separate management layer |
Control over tasks and graders | Total | Through the spec and acceptance criteria | Total; the work happens in your repository |
Where the code lives | Your repository | The vendor's platform or a delivered package | Your repository |
What you own at the end | Everything, plus the people who built it | The delivered artifact, on the licence you negotiated | The code, the harness, the graders and the documentation |
What breaks when the model changes | Your team rebuilds on its own schedule | You buy the next version or a refresh | The engineers who built it extend it |
Fits when | The environment is the research, or the domain is one your engineers practise | You need breadth across domains, or a dataset fast | The work is engineering you need to keep extending |
Strains when | The hiring cycle is longer than the roadmap | The workflow is specific to you and keeps moving | You need dozens of domain experts, not engineers |
When does building in-house win?
Build when the environment is part of the research question rather than an input to it. If your team is studying how reward design changes what a model learns, the grader and the task distribution are the experiment, and handing them to anyone outside the research loop slows the loop down. Build also when the domain is one your engineers already practise. Invisible's own argument about why labs outsource concedes the point: coding and maths stay in-house because the ML engineers are themselves practitioners and know what a correct answer looks like.
The cost of building is time, then load. A senior engineer who has built rollout infrastructure and adversarial graders before is a hard hire, and every seat runs its own search. Once they're in, environment work is lumpy: a burst when a new environment family stands up, a long tail of maintenance, another burst at the next target. Teams that size permanent headcount to the burst carry idle seats; teams that size it to the tail wait on hiring every time. The staffing guide for RL environments work covers the usual answer, a small permanent core that holds the library and the harness plus a flexible senior layer around each push. Building also concentrates knowledge: if two engineers hold the whole harness in their heads, the library has a bus factor of two.
When does buying finished tasks or environments win?
Buy when you need breadth faster than any team could build it. A model that needs to handle legal review, insurance claims, and procurement workflows in one quarter needs domain experts in all three, and a vendor with a large expert pool and a working quality process can staff that in a way your hiring plan can't. That's the strongest version of the argument for buying, and it's correct.
Buy also when someone has already built the environment. A catalogue replica of a common SaaS tool or a dataset of coding tasks in real repositories exists, and paying to rebuild it is waste. Some vendors now price around the result. Surge's off-the-shelf programme lets labs in its Trusted Program train on a dataset, evaluate the model, and pay if it moves the metrics they care about, which moves part of the risk onto the vendor.
What you're buying is an artifact on the vendor's schedule, under the vendor's process. The spec and the acceptance criteria are your control surface, so they need to be precise about the end states the grader checks, what counts as a pass, and how the vendor tested the grader against an agent trying to cheat. The quality question doesn't go away when you buy. One interviewee in Epoch's FAQ put the vendor's problem plainly: "Finding the experts isn't that hard, but managing them and doing quality control is hard." That work happens inside the vendor, out of your sight, and the result reaches you as a finished package.
Ask early what the licence covers. Whether the vendor can sell the same environment to the lab down the road, whether you get the grader source or only the scores, and whether you can modify the environment yourself are contract questions with very different answers across vendors.
When does staffing senior engineers under your roadmap win?
Staff when the thing you need is engineering you'll keep extending: the replica of your enterprise app, the seeded data, the tool layer, the graders that check end state, the harness that runs rollouts and stores trajectories. Those are long-lived systems. They change every time the model changes, and the people who built them are the cheapest people to change them.
Staffed capacity sits between the other two routes. Your researchers set the roadmap and decide what the environments need to teach. Senior engineers commit into your repository under your review. A senior lead arrives inside the team, in a builder seat, owning the plan, the quality bar, and coordination across the group, so your research leads aren't spending their week running a contractor pool. There's no management layer priced above the team. That's the embedded-lead model in the comparison of individual contractors and managed teams, and it's the right one for environment work because the lead is close enough to the code to catch a gamed grader before a researcher does.
What you keep is the system and its documentation, in your repo. What makes the model work across more than one push is continuity. If the same named engineers come back for the next environment family, the second engagement starts without a ramp. If the vendor reshuffles its pool between engagements, you pay the ramp every time, and the reason grader v4 exists walks out with whoever wrote it.
Staffing has limits worth stating. It doesn't give you two hundred domain experts across a dozen professions; that's a marketplace's strength, and for breadth of expert judgement a marketplace is the right buy. It doesn't remove the need for someone on your side who decides what the environments are for. And it only beats hiring if the engineers are genuinely senior, which is why you should interview the people you'd get by name before anyone starts.
A.Team sits in this column: senior engineers matched to environments work, arriving with a lead inside the team and no managing-partner fee, committing into your repository under your researchers' direction.
What breaks when the model changes?
Three things break, and they break on every checkpoint. Tasks saturate, so the model passes most of them and the signal disappears. Graders get beaten in ways the last checkpoint never tried. And the harness needs to cover tool surfaces and longer horizons the old one didn't.
Saturation is the quiet one. A task set calibrated against last quarter's model keeps running after the new checkpoint clears it, and the numbers stop telling you anything. Someone has to write harder tasks, retune the difficulty spread, and retire the tasks that no longer separate good from bad.
Grader failures are the loud one. A more capable model finds routes to a pass that a weaker one didn't: editing the check, writing the expected output directly, satisfying a proxy instead of the task. The staffing guide lays out the defences. What matters here is who applies them when the new failure shows up.
Under each route that answer is different. If you built in-house, your team fixes it on its own schedule, which is fine if the team still exists and has time. If you bought, you wait for the vendor's refresh or commission a new version, and you depend on the vendor having kept the people who understood your environment. If you staffed, the engineers who wrote the grader are the ones who patch it, in your repo, the same week. That's the row of the table to weigh most heavily for enterprise workflows, because those workflows change on their own as well, every time the business changes a process or a SaaS vendor ships an update.
Are there cost-effective RL environments for startups?
Yes, if cost-effective means paying only for the part you can't get elsewhere. For a startup, that usually means starting from open environments, buying catalogue items where they fit, and paying for engineering only on the custom layer that encodes your workflow.
The open end is larger than it was a year ago. Prime Intellect's Environments Hub lists 2,500+ community environments, alongside its open-source Verifiers library for building environments and its prime-rl training framework. Starting there means your engineers spend their hours on the custom layer and inherit the plumbing.
After that, the unit of sale matters more than the headline price. Per-task pricing is cost-effective when the tasks are stable and you need many of them. Per-environment pricing works for a catalogue replica you'll use as delivered. Per-engineer pricing works for the custom layer that has to keep changing, because the cost of a change is an engineer's afternoon rather than a new statement of work. A startup that pays per task for work that changes weekly pays for the same environment several times over.
The trap is buying a finished custom environment for a workflow that's still moving. If your product's own workflow is a quarter old, the environment that encodes it will be out of date before the training run ends.
What to do next
List every environment on your roadmap for the next two quarters and mark each one on two axes: generic or specific to your workflow, and stable or changing. Generic and stable goes to a catalogue vendor or an open environment. Specific and stable can be bought as a custom build if the licence gives you the grader source. Specific and changing is engineering, and it belongs either on your payroll or with a staffed team working in your repo. Then compare vendors on the unit of sale each one actually sells, and if you're replacing an expert-hours vendor, read the alternatives to Mercor and Scale when the work is engineering.
For the specific-and-changing column, A.Team puts senior engineers for environments work in front of you, with a matched shortlist within 72 hours of the scoping call, from a network of 11,000+ vetted builders with under 2% acceptance. On our largest environments engagement, at an RL-environments and agent-training infrastructure company, 45 senior builders were ramped and running inside a week, and 193 started across the engagement. Many of them grew into pod leads running their own teams.
Frequently asked questions
Where custom RL environments come from, what vendors sell, how they're built and who owns the IP.
Three places. Vendors such as Scale and Turing sell finished enterprise environments, including replicas of common SaaS tools with verifiers attached. Vendors such as Invisible build custom environments as a service. Or senior engineers, on your payroll or staffed under your roadmap, build them in your repository. Catalogue items fit generic workflows; workflows specific to your business that keep changing are usually better built by engineers you keep.
Usually some of each. Start from open environments and shared verifier libraries where they exist, buy catalogue datasets or environments for generic workflows, and spend engineering time only on the custom layer that encodes your own workflow. Buying a finished custom environment for a workflow that's still changing tends to cost more over a year than staffing the engineering, because every change becomes a new order.
Four different things, which are easy to confuse: expert-hours from a contributor marketplace, finished datasets and environments, platforms and tooling you run yourself, and engineering capacity that works in your codebase. Many companies sell more than one. Sorting vendors by unit of sale before comparing them on quality or price saves a lot of mismatched calls.
You need a reproducible starting state for each rollout, a task specification, tools or an interface the agent acts through, and a grader that checks the end state and can't be satisfied without doing the task. Around that sits a harness that runs rollouts at scale and stores trajectories. For enterprise workflows the starting state is usually a replica of the app, seeded with realistic data.
It depends entirely on the contract. Some vendors deliver environments under a licence that lets them sell similar environments to other labs, some give you scores but not grader source, and some assign everything to you. When engineers build in your repository under your direction, the code sits with you from the first commit. Ask the licence question before the scoping call.

How to staff the engineering under an RL roadmap
Most of the work under an RL or agent-training roadmap is senior software engineering: sandboxed execution, job orchestration, graders that survive an adversarial agent, and eval pipelines a researcher will actually trust. The vendors that dominate this market sell task volume out of contributor pools on three-month clocks, which is a different product from engineering capacity.

RL environment companies in 2026: Sorted by what they actually sell
Top RL environment providers in 2026, sorted by unit of sale: expert-hours, finished datasets and environments, platforms, and engineering capacity.

Alternatives to Mercor and Scale when the work is engineering
Scale AI alternatives and Mercor alternatives, sorted by what each sells: expert-hours, finished data, platforms, or engineering capacity.
Staff the engineering under your RL roadmap
A.Team matches senior, AI-native engineers who design, build, and validate the environments, tasks, and evals your models train on, with a lead who owns delivery from inside the team. On our largest environments engagement, 45 senior builders were ramped and running inside a week. Tell us what you're training and we'll match builders in days.
Scope Your Environments Work