Research planning

Gameplay data for world models and imitation learning

A recording can show what happened, what a player did, or both. The right gameplay dataset depends on which relationship you want a model to learn. Start with the training target, then decide whether you need recorded actions, video alone, or a combination.

World models learn how observations change

An action-conditioned visual world model uses past observations and actions to predict how a scene evolves. Relevant data includes the visual sequence, the action representation, and a clear convention connecting them. Continuity matters because a transition should reflect the environment responding to play, rather than an unmarked edit.

GameNGen illustrates this approach: the researchers recorded an agent’s play and trained a diffusion model using previous frames and actions to predict the next frame. Its data source was an RL agent. That result demonstrates one research pipeline; it does not establish that human recordings are required for all world models.

For a human gameplay collection, specify situations your model needs to represent: navigation, object interaction, recovery from mistakes, changing viewpoints, and transitions between areas. Ask whether the retained sequences preserve those events and whether their action labels match the interface your model will use.

Some world-model approaches learn from video alone

Action labels are not a universal prerequisite. The original Genie research learned an interactive environment from unlabeled videos using a latent action model. Its learned action representation is different from directly recording a person’s keyboard and mouse.

The practical implication is to specify the representation you need. If your method learns latent actions, video continuity and coverage may be your main collection concerns. If you need an interface tied to actual keys or measured mouse movement, ask for those labels explicitly. Do not assume a learned latent action has a ready-made correspondence to a game’s controls.

Imitation learning needs a behavior to imitate

In supervised behavior cloning, a policy learns to predict demonstrated actions from observations. That makes the definition and quality of those demonstrations central to the collection brief. “Expert gameplay” is too broad: an experienced competitive player, a careful explorer, and a player following a specific task may produce different training signals.

Decide whether the policy should learn successful execution, recovery, exploration, or a mixture. Request task instructions, outcome labels, or player experience categories if your experiment depends on them. The current collector does not automatically infer those annotations from a recording.

Preserve enough history for the decision. A single image may omit the goal, an earlier instruction, or information seen elsewhere in the environment. Define your observation window and evaluate whether the demonstration package can support it.

Inverse dynamics can connect labeled and unlabeled data

Video PreTraining (VPT) used labeled demonstrations to learn an inverse dynamics model, then applied it to unlabeled Minecraft video to generate action labels for policy training. This separates two collection needs: data for learning how visible changes relate to actions, and video covering behaviors the policy should learn.

If you use a similar strategy, retain label provenance. A recorded input and an action predicted by another model should not silently become the same category in your manifest. Evaluate inferred labels on held-out recordings with actual inputs, and inspect errors for the actions and settings that matter to your task.

This is an experimental option, not a claim that labels inferred in one title will transfer to another. Camera perspective, interfaces, visual style, and control settings can differ substantially.

Turn the approach into collection requirements

Training approachData to specifyA useful acceptance question
Action-conditioned world modelObservation sequences, actions, timing convention, interruption policyCan I build valid next-observation examples across the sequences I need?
Video-only or latent-action modelVideo coverage, sequence length, visual settings, continuityDoes the footage preserve the temporal changes my method learns from?
Behavior cloningDemonstrations, recorded actions, task definition, relevant contextDoes the demonstrated behavior match what I want the policy to do?
Inverse dynamicsPaired video/actions, label provenance, held-out recorded inputsCan I measure action prediction error on the intended distribution?

These are starting points for a brief, not mutually exclusive product categories. One collection can support several experiments if its coverage and annotations are designed for them.

Use a pilot to test the data assumptions

Build a small end-to-end data path first: decode observations, map controls, handle interrupted sequences, and run your evaluation split. Keep sessions together when testing generalization to new sessions; use player, map, or title separation when those are the claims being tested.

Gaming Datasets’ current Windows collector produces gameplay video with keyboard and mouse sidecars. It does not expose privileged engine state, rewards, calibrated camera angles, or automatic success labels. Review the action format and its limits, then use the dataset evaluation checklist to prepare a pilot.

Tell us your training objective and collection requirements. We can discuss feasibility and the fields to agree before collection. The research cited here provides methodological examples and does not imply a relationship with those research teams.

These guides explain collection decisions and the current recorder format. Dataset availability, game compatibility and delivery requirements are agreed per project. For corrections or a technical question, contact Gaming Datasets.