Data Products

License game trajectories. Build RL environments. Collect Asia-scenario data. Deliver.

PLAYER TRAJECTORIESRL ENVIRONMENTSAGENT TRACESEGOCENTRIC VIDEODIALECT SPEECHDOCUMENTS & E-COMMERCE

Data You Can't Scrape

Frontier labs have exhausted the open web. What's left — real trajectories with actions, state, and outcomes — has to be licensed, built, or collected. Public supply of the categories that matter most is near zero: Asian-app agent trajectories, iOS traces beyond eval sets, long-horizon cross-application workflows, and fresh uncontaminated eval data. That's what we produce.

Commercially Licensed

The strongest open datasets are research-only or CC-NC — fine to benchmark on, illegal to ship a product trained on. Ours arrive consent-cleared, with model releases, source authorization, and a rights report every time.

Schema-Complete

Video, inputs, engine state, task goals, and outcomes arrive frame-synchronized in one schema — no reverse-engineering signals from raw footage.

Continuously Generated

Self-built environments and standing studio integrations mean data keeps flowing after delivery — not a one-time archive dump.

Asia-Native

WPS, LINE, KakaoTalk, e-commerce back ends, local languages, and regional scenarios — coverage that Western collection pipelines can't reach.

20–50 tasks
50–200 trajectories
1 data schema
1 quality & licensing report

How Engagements Start

Every engagement starts as a paid pilot: a small, priced batch that proves task design, schema, quality, and licensing before you commit to scale.

Game Data Network

Asia's Games, Licensed for Training

Player trajectories from live titles, indie studios, and sunset projects that still hold their source code.

Trajectory Viewer
task: "clear warehouse level 3" · result: success · 4m12s

Video · Inputs · Engine State · Outcome

Licensed Player Trajectories

Sourced, cleared, and captured at the engine level

We screen Asian studios and sunset titles, run rights diligence, negotiate AI-training licenses, and integrate directly with the game engine — so every trajectory arrives with video, inputs, state, task goals, and outcomes in sync.

Agent Trajectories

The Full Stack, Not Just Screenshots

Public agent data rarely pairs more than screenshots and actions. We capture every layer of every step — the difference between data a model can imitate and data a model can reason from.

Per-Step Schema

  • Screenshot
  • Accessibility tree
  • DOM snapshot
  • Screen recording
  • Typed action — executable code, normalized coordinates
  • Step-level reasoning annotation
  • Failure-and-recovery traces

Coverage the public sets don't have

Where public supply runs dry, we produce to spec: agent trajectories inside Asian super-apps and domestic enterprise software; iOS traces beyond eval-only sets; long-horizon 50–200-step workflows that cross applications, where public sets average around twenty steps; and fresh, time-stamped eval sets built after your model's cutoff so scores aren't contaminated by training data.

Enterprise Workflow Data

How Real Work Actually Moves

Models trained on the public web have never seen how work travels through a company. We partner with operating businesses to extract anonymized workflow data from the production tools their teams use every day, so agents can learn the patterns that exist only inside real enterprise software. We deliver this data through two paths, so you can choose the provenance profile that fits your program, or combine both.

Path A · Anonymized Extraction From Real Companies

Signals We Deliver

  • Message thread activity across chat and email
  • Document creation and edit histories
  • Meeting records and shared notes
  • Task, ticket, and project completion
  • Calendar and scheduling events
  • CRM pipeline and code repository activity

Anonymized before it ever leaves the source

Contributing companies connect their tools through read-only access for a single extraction, and every name, address, credential, and sensitive field passes through an automated pipeline that masks it irreversibly at the source. You receive the structure of real work with no identities attached, across messaging, email, documents, project management, CRM, code repositories, and finance.

Path B · Expert-Recorded Trajectories in High-Fidelity Sandboxes

Inside Every Trajectory

  • Step-by-step actions across every tool
  • Tool switches and context changes
  • Intermediate drafts and working artifacts
  • Corrections and revision history
  • The final deliverable for each task
  • Grading rubric and reference answer

Demonstrated by professionals, recorded end to end

We recruit practitioners with years of real tenure in fields such as investment banking, consulting, accounting, customer support, and software engineering, then place them in high-fidelity sandbox replicas of the tools their jobs run on, from collaboration suites to CRM and finance systems. They complete designed work tasks while we record the full trajectory, so your model learns how the work gets done rather than only what the answer looks like. Structured for SFT, evaluation, and agentic RL. Every contributor signs an IP assignment and a data release, no company data is ever touched, and we can warrant the provenance of every record in your contract.

Simulation & RL Environments

Environments With Verifiable Reward

The market has moved past static SFT dumps: what trains a frontier agent now is an executable environment where reward is checked by execution, not opinion. Ours are resettable, parameterized, and programmatically auto-graded — with interactive objects, explicit task goals, and full engine-state output — starting with indoor navigation, warehouse logistics, construction, city driving, and Asian computer use. Sold as static trajectories, task generators, or executable RL environments.

Production Pipeline

Sourcing

Screen Asian studios, sunset titles, and contributor pools for the signals your model needs.

Licensing

Run rights diligence and negotiate AI-training licenses with a documented chain of title.

Capture

Integrate with the engine or device to record trajectories: screen, inputs, state, tasks, outcomes.

QC & Delivery

Clean, verify, and deliver against an agreed schema, with a quality and licensing report.

Asia-Scenario Collection

What We Collect On Demand

Game Trajectories

  • Gameplay Video & Inputs
  • Engine State Sync
  • Task Goals & Outcomes
  • Success / Failure / Recovery

Agent Traces

  • Asian Desktop (WPS, Feishu, LINE)
  • Screenshot · A11y Tree · DOM Snapshot
  • Typed Actions, Normalized Coordinates
  • Step Reasoning · Failure & Recovery

Egocentric Video

  • Packing & Sorting
  • Assembly & Tool Use
  • Search & Carry Tasks

Speech & Language

  • Consent-Cleared Asian Languages
  • Mandarin · Japanese · Korean · Thai · Vietnamese
  • Spontaneous Conversation, Not Just Read
  • Asian Documents & E-commerce

Expert Reasoning Data

Problems Frontier Models Fail

Open reasoning datasets are overwhelmingly model-distilled and benchmark-contaminated. Ours are written from scratch by a network of domain experts, verified, and rubric-graded — and both our clients and our contributors know exactly how the data will be used.

Original & Uncontaminated

Graduate students, PhDs, and practitioners across medicine, law, finance, and STEM author new problems that frontier models still get wrong — not scraped, not distilled.

Verified & Rubric-Graded

Every item ships with a verifier and a grading rubric, with step-level process supervision for PRM-ready reward modeling.

Asian Languages & English

Multilingual authoring across every domain, so your model learns to reason in each language rather than translate after the fact.

Expert Produced Data

Work Products From Working Professionals

Some capabilities cannot be scraped, because the knowledge lives in practitioners' heads and surfaces only in the work they produce. We contract active and former professionals across investment banking, consulting, law, accounting, and software engineering to create the artifacts their jobs actually generate, turning tacit expertise into training data.

Real Deliverables

Industry reports, analysis memos, task walkthroughs, evaluation scores, and gold-standard answers, authored by people who do the work rather than crowd workers approximating it.

Built for Every Stage

Structured for supervised fine-tuning, held-out evaluation, and reinforcement learning, with grading rubrics and reference solutions on request.

Vetted for Real Tenure

Every contributor is screened for active or prior experience in the field they write for, so the reasoning reflects how the job is genuinely done.

“Datasets you can scrape, someone else already has. The trajectories that move a frontier model — real actions, real state, real outcomes, with a clean license — have to be produced. We built our supply chain across Asia to produce exactly that.”

Why we exist

Emerge

Founding Team

Your model's weakest task is our starting point