Commercially Licensed
The strongest open datasets are research-only or CC-NC — fine to benchmark on, illegal to ship a product trained on. Ours arrive consent-cleared, with model releases, source authorization, and a rights report every time.

License game trajectories. Build RL environments. Collect Asia-scenario data. Deliver.
Frontier labs have exhausted the open web. What's left — real trajectories with actions, state, and outcomes — has to be licensed, built, or collected. Public supply of the categories that matter most is near zero: Asian-app agent trajectories, iOS traces beyond eval sets, long-horizon cross-application workflows, and fresh uncontaminated eval data. That's what we produce.

The strongest open datasets are research-only or CC-NC — fine to benchmark on, illegal to ship a product trained on. Ours arrive consent-cleared, with model releases, source authorization, and a rights report every time.
Video, inputs, engine state, task goals, and outcomes arrive frame-synchronized in one schema — no reverse-engineering signals from raw footage.
Self-built environments and standing studio integrations mean data keeps flowing after delivery — not a one-time archive dump.
WPS, LINE, KakaoTalk, e-commerce back ends, local languages, and regional scenarios — coverage that Western collection pipelines can't reach.
20–50 tasks
50–200 trajectories
1 data schema
1 quality & licensing report
How Engagements Start
Game Data Network
Player trajectories from live titles, indie studios, and sunset projects that still hold their source code.
Video · Inputs · Engine State · Outcome
Licensed Player Trajectories
Sourced, cleared, and captured at the engine level
We screen Asian studios and sunset titles, run rights diligence, negotiate AI-training licenses, and integrate directly with the game engine — so every trajectory arrives with video, inputs, state, task goals, and outcomes in sync.
Agent Trajectories
Public agent data rarely pairs more than screenshots and actions. We capture every layer of every step — the difference between data a model can imitate and data a model can reason from.
Per-Step Schema
Coverage the public sets don't have
Where public supply runs dry, we produce to spec: agent trajectories inside Asian super-apps and domestic enterprise software; iOS traces beyond eval-only sets; long-horizon 50–200-step workflows that cross applications, where public sets average around twenty steps; and fresh, time-stamped eval sets built after your model's cutoff so scores aren't contaminated by training data.
Enterprise Workflow Data
Models trained on the public web have never seen how work travels through a company. We partner with operating businesses to extract anonymized workflow data from the production tools their teams use every day, so agents can learn the patterns that exist only inside real enterprise software. We deliver this data through two paths, so you can choose the provenance profile that fits your program, or combine both.
Path A · Anonymized Extraction From Real Companies
Signals We Deliver
Anonymized before it ever leaves the source
Contributing companies connect their tools through read-only access for a single extraction, and every name, address, credential, and sensitive field passes through an automated pipeline that masks it irreversibly at the source. You receive the structure of real work with no identities attached, across messaging, email, documents, project management, CRM, code repositories, and finance.
Path B · Expert-Recorded Trajectories in High-Fidelity Sandboxes
Inside Every Trajectory
Demonstrated by professionals, recorded end to end
We recruit practitioners with years of real tenure in fields such as investment banking, consulting, accounting, customer support, and software engineering, then place them in high-fidelity sandbox replicas of the tools their jobs run on, from collaboration suites to CRM and finance systems. They complete designed work tasks while we record the full trajectory, so your model learns how the work gets done rather than only what the answer looks like. Structured for SFT, evaluation, and agentic RL. Every contributor signs an IP assignment and a data release, no company data is ever touched, and we can warrant the provenance of every record in your contract.
Simulation & RL Environments
The market has moved past static SFT dumps: what trains a frontier agent now is an executable environment where reward is checked by execution, not opinion. Ours are resettable, parameterized, and programmatically auto-graded — with interactive objects, explicit task goals, and full engine-state output — starting with indoor navigation, warehouse logistics, construction, city driving, and Asian computer use. Sold as static trajectories, task generators, or executable RL environments.

Production Pipeline
Screen Asian studios, sunset titles, and contributor pools for the signals your model needs.
Run rights diligence and negotiate AI-training licenses with a documented chain of title.
Integrate with the engine or device to record trajectories: screen, inputs, state, tasks, outcomes.
Clean, verify, and deliver against an agreed schema, with a quality and licensing report.
Asia-Scenario Collection
Expert Reasoning Data
Open reasoning datasets are overwhelmingly model-distilled and benchmark-contaminated. Ours are written from scratch by a network of domain experts, verified, and rubric-graded — and both our clients and our contributors know exactly how the data will be used.
Graduate students, PhDs, and practitioners across medicine, law, finance, and STEM author new problems that frontier models still get wrong — not scraped, not distilled.
Every item ships with a verifier and a grading rubric, with step-level process supervision for PRM-ready reward modeling.
Multilingual authoring across every domain, so your model learns to reason in each language rather than translate after the fact.
Expert Produced Data
Some capabilities cannot be scraped, because the knowledge lives in practitioners' heads and surfaces only in the work they produce. We contract active and former professionals across investment banking, consulting, law, accounting, and software engineering to create the artifacts their jobs actually generate, turning tacit expertise into training data.
Industry reports, analysis memos, task walkthroughs, evaluation scores, and gold-standard answers, authored by people who do the work rather than crowd workers approximating it.
Structured for supervised fine-tuning, held-out evaluation, and reinforcement learning, with grading rubrics and reference solutions on request.
Every contributor is screened for active or prior experience in the field they write for, so the reasoning reflects how the job is genuinely done.
Resources
“Datasets you can scrape, someone else already has. The trajectories that move a frontier model — real actions, real state, real outcomes, with a clean license — have to be produced. We built our supply chain across Asia to produce exactly that.”
Why we exist
Emerge
Founding Team