Reward Arcadea hands-on lab

What an RL environment actually is

Policy Pond taught the learner and Finishing School taught the LLM learner. Neither taught the world they practice in. This is that world: a sealed cabinet with three controls, a ticket counter inside, and a strict rule about what the player may see. About 75 minutes of play, with a real environment running in your browser.

0 · The cabinet: three controls, one game per seat

An environment is the world an agent practices in. For an LLM, OpenEnv packages that world as a sealed box running in its own container, with exactly three controls. Every environment has the same three, whether it plays Hangman, Wordle, a code repo or a fake CRM. Every term with a wavy red underline is tappable: hover or tap it for the plain meaning and its arcade equivalent.

The environment = the cabinet
a sealed box; the game logic lives inside
reset() = the coin slot
a fresh game; returns what you see first
step(action) = the joystick
one move; returns the screen, the tickets, and whether it's over
Observation = the front glass
everything the player sees
state() = the back panel
the cabinet's own bookkeeping, behind a service hatch
Reward = tickets
paid per move by the cabinet, never by the trainer
done = the game-over light
one flag; it doesn't say why the game ended
Policy = the player
a function from what it sees to a move; later, an LLM
Eval = the tournament scoreboard
the same cabinet, scores read by people

Covers OpenEnv course Module 1 · the RL loop, the 3-method interface · Module 2 · typed models

Play the cabinet

This is the course's own Module 4 game, running here logic for logic. A hidden six-letter word; one letter per move; ten wrong guesses and you lose. Drop a coin, then press letters. The tape below prints exactly what each call returned.

How you're talking to the cabinet
0
step_count
0
tickets this game
no
done
0
games your presses touched
The ticket tape: what each call returned
Guess first!
Switch to one-off requests and press a letter, then another. What does the second press see?
A brand-new game. OpenEnv's plain HTTP /reset and /step are stateless: each call builds a fresh environment, answers, and throws it away. Only a WebSocket session keeps one game alive between calls, which is why the typed client uses /ws. Verified on the real server: two one-off steps of the same wrong letter both reported 9 attempts left.
Guess first!
Two trainers open two connections to one cabinet server. Do they share a game?
No. Each WebSocket connection gets its own environment instance on the server. That's how one container can serve a whole training batch at once. One catch, from TRL's docs: an OpenEnv server allows 1 concurrent session unless its app.py raises max_concurrent_envs, so a trainer that opens 64 connections to a default server fails. Set it to at least the generation batch size.
Say it out loud
Reset drops the coin. Step moves the joystick once and returns what I see, the tickets, and whether it's over. State is the back panel. One connection is one game.

1 · Front glass, back panel: who may see what

The most important design decision in any environment is made in its types file: which fields go on the observation (the front glass, what the player sees), which go in the state (the back panel: bookkeeping the client can still fetch), and which never leave the server. In RL terms the state is the world's full true situation and the observation is the player's partial view of it. The course's word game puts the secret word in state. Sort the fields yourself, then let a cheater play against your layout.

Covers Module 4 · Step 1, models.py · Module 1 · observation vs state · the study pack's M4 gotcha 5

Sort the eight fields

Tap a field, then tap a bin. (On a computer you can also drag.) The cabinet needs the word to run the game; the question is where it may sit.

Front glassthe observation: returned by every step()
Back panelthe State model: returned by state()
Sealed insidea private attribute on the server object; never sent
what env.state() returns to the client
what the player sees after a move
against Random, on your layout
–
Cheater's win rate
–
Random's win rate (the floor)
Guess first!
The course keeps target_word in the State model. Can a player read it?
Yes. env.state() is one of the three client methods, and the server answers it with the whole State model. Verified against the real server: state() came back with target_word='tensor' mid-game. A training loop that ever puts state() in the prompt, or an agent with a tool that calls it, learns to read the answer instead of playing.
Guess first!
You "fix" it by deleting target_word in the client's _parse_state. Is the leak closed?
No. The bytes still cross the wire; anyone can write their own client from the server's /schema in twenty lines. A secret has to stay on the server, as a private attribute the game logic reads and nothing serializes. Fix the server's reset(), not the client.
Say it out loud
Anything the client can fetch, assume the player can see. Secrets stay on the server.

2 · Shake the joystick: inputs a good player never sends

An optimizer doesn't play politely. It sends whatever pays. So before anything trains against a cabinet, send it the moves a good player never would and read what comes back. The course's game has four such holes, all of them found by running it. Each switch below closes one.

Covers Module 4 · Step 2, environment.py · TRL's "raise on invalid actions" rule · the study pack's M4 gotchas 4, 8 and 9

The cabinet's switches (all off = the course as written)
Send an odd input
What came back
The four holes
Guess first!
Press step before a coin on the course's cabinet. What does it pay?
A win. A fresh cabinet's word is the empty string; masking it leaves no blanks, so "no blanks left" reads as solved: reward 1.0, done, and the message "You got it! The word was ''." Over plain HTTP, where every call is a fresh environment, every step is a win. A rollout that ever skips reset() scores 100% for nothing.
Guess first!
Leave the step cap off and press the same right letter ten times. What can a player do to the course's cabinet?
Play forever. In the course's game a repeated correct letter costs no attempt (only wrong guesses do), and the environment has no step limit, so a player pressing the same right letter never reaches game over. The notebook's guard doesn't change that; it only stops a repeated wrong letter from being charged twice. Harmless, until step 3 hangs a prize on it. Every cabinet needs a clock.
Say it out loud
Check what the cabinet does on inputs a good player would never send. An optimizer will send them. Errors pay nothing; that's the safe answer.

3 · The prize rules: which player would training pick?

A trainer doesn't know what you meant. It raises the average tickets. So for any reward rule the only question is: which player does it rank first? If that isn't the player you want, training will find the one you don't. Three rules, three players. The Farmer never wins a game; it finds one correct letter and presses it forever.

Covers Module 5 · reward functions, shaped vs binary · TRL's OpenEnv docs, "Tips for reward functions" · Finishing School step 6 (the group is the critic)

The prize rule (set on the cabinet, because the environment owns the reward)
Step cap (the harness's clock)50
Average tickets per game · Frequency and Farmer are exact (one game per word, deterministic) · Random is simulated over 300 games

The group of 2: what GRPO sees

GRPO scores a group of games on the same prompt against their own average. The course trains with groups of 2. If the two games score the same, the advantage is zero and that group teaches nothing. Same player, same word, two games: how often do they score differently under the rule above?

Guess first!
Rule 2 pays +0.1 every time the guessed letter is in the word. Who scores highest?
The Farmer, at about 4.9 tickets a game against Frequency's 0.67. It never wins; it collects 0.1 on every one of ~49 presses of the same letter, because rule 2 pays for touching a correct letter, not for progress, and repeats are free of consequence. TRL's own experiments found binary rewards "gave cleaner training signals than shaped rewards" for exactly this reason. Rule 3 pays only for new letters and caps the total, so the Farmer banks one letter's share (~0.09) and stops earning.
Guess first!
Frequency presses the same letters in the same order every game. What does a GRPO group of two Frequency games teach?
Nothing. Two identical games score identically, whatever the word, so the group has no spread and every advantage is 0. GRPO needs the player to vary: for an LLM that's sampling temperature; a greedy decoder would have the Farmer's problem in reverse. And under the binary rule even Random's two games differ only ~5% of the time, since both usually lose; under rule 3 they differ ~79% of the time. That's the honest case for shaping when groups are tiny and wins are rare.
Say it out loud
The trainer maximizes the tickets, whatever I meant. Pay for progress, cap it, and make sure the player varies, or the group teaches nothing.

4 · The tournament: the same cabinet, read by people

An eval is the cabinet with a scoreboard: a fixed set of games, a verifier inside, and a pass rate for a human to read. Module 2's "policy competition" was one, without the two things that make a score believable: a floor and a ceiling, and error bars. This scoreboard has both, plus an A/A test: two identical Random players who must post the same score, or the harness itself is broken.

Covers Module 2 · policies and the competition · the LangChain evals vocabulary (verifier, environment, harness) · TRL's "if a capable model can't beat random…"

Games per player250
Frequency's true win rate is exactly 20%: it wins on 2 of the 10 words (neural, tensor) and no others. The dashed line marks it. Watch how far a single run lands from the truth.
Frequency, ten separate runs at this size (true rate 20%)
Guess first!
Frequency's true rate is exactly 20%. Over 100 games, how far off can one run land?
Seven points or more, easily. The standard error of a 20% rate over 100 games is √(0.2 × 0.8 / 100) = 4 points, so one run in three lands more than 4 points off and one in twenty more than 8. When the lab behind this page was built, ten runs of 100 games read anywhere from 14% to 28%. One eval run is a sample.
Guess first!
Random A scores 5% and Random B scores 2% over 250 games each. Is the harness biased?
No. Pooled over 500 games at ~3.5%, the standard error of the difference is about 1.6 points, so a 3-point gap is under two standard errors: ordinary luck. The test only alarms past 3.5 standard errors, where a correct harness false-alarms about once in ten thousand runs. A first draft of this test alarmed 1.7% of the time on a working harness, which is how often you'd have "found a bug" that wasn't there.
Say it out loud
One run is a sample. Floor, ceiling, error bars, and an A/A test before I believe a gap. A hundred percent means somebody read the answer.

5 · Arcade vs pond: when to put the cabinet on the network

The course argues that in-process environments (Gym-style) are "a disaster for production." That's true for LLM training, where every step waits ~200 ms for the model to think and the game itself may be a code sandbox you don't want inside your trainer. It's the opposite for Microduck, whose 4,096 ducklings advance together as one tensor on one GPU; a network hop per step would cost more than the whole step. Move the sliders and watch which term owns the wall clock.

Covers Module 1 · why Gym falls short · the course's scaling appendix (WebSocket vs HTTP) · Policy Pond step 6 (Microduck's real config)

Preset
Policy time per step (the player thinking)
Environment time per step (the game logic)
Network round trip per step
Environments stepping together
Steps per rollout
A teaching model. The times are order-of-magnitude defaults on sliders, not measurements; the arithmetic (one step's wall time = policy + environment + network) is the whole model.
Guess first!
On the Microduck preset, what share of a service-style step is the network?
Nearly all of it: about 99%. A batched physics tick and a small MLP forward pass together take a fraction of a millisecond for all 4,096 ducks, and a network round trip is tens of milliseconds. That's why mjlab keeps every duck in the trainer's process as GPU tensors, and why OpenEnv doesn't. Same loop, opposite economics.
Say it out loud
Put the cabinet on the network when the player is slow or the game is dangerous. Keep it in-process when the game is a physics tick.

★ · Review board: three broken cabinets

Three environments came back from training with suspicious scores. For each, read the symptom and the evidence, name the cause, then pick the fix. The tempting wrong fixes are there too, and each fails for a stated reason.

Covers everything above · the study pack's evals bridge §4 checklist

Guess first!
A model scores 100% on a public benchmark whose problems and solutions are on GitHub. What did you learn about its skill?
Less than it seems. That's a leak at the dataset level, called contamination: the answer sheet was in the training data rather than in the state. A private task the model has never seen is the cleaner test. The word "leakage" covers both.
Say it out loud
A hundred percent is a bug report. Find where the answer leaked or what the reward paid for, and fix the cabinet, not the player.

✓ · Field test

Eight questions, each answered by setting up a widget in steps 0–5 the way the question says and reading off the result. Graded instantly, with rounding tolerance.

Fix all three cabinets on the review board and score 6 or better here, and the arcade posts your high score.

What's exact and what's a teaching model

Exact: the Hangman environment is the OpenEnv course's Module 4 game, logic for logic, including its bugs, and its behaviour on the real server (openenv-core 0.3.0) was checked call by call before this page was written. Frequency's and the Farmer's averages are exact expectations: one game per word over the course's ten words, drawn uniformly. The standard errors and the A/A test are the plain binomial formulas.

Simulated: Random's and the Cheater's scores, and the "two games differ" rates, come from a seeded pseudo-random generator, so they repeat exactly across reloads but are samples, not truths (Random's true win rate is about 2.3%). Illustrative: step 5's times are order-of-magnitude defaults on sliders; the review board's three cases are composites, not incidents.

The environment is real; the numbers about it were measured; the timing model is a sketch.

Where to go deeper

  • The OpenEnv course (five modules with notebooks), and the OpenEnv repo: its rfcs/ folder explains rubrics (composable rewards) and agentic harnesses.
  • TRL's OpenEnv integration: environment_factory, the reward-function tips this page quotes, and the concurrency note.
  • Policy Pond for the learner (REINFORCE, PPO, GAE) and Finishing School for the LLM learner (SFT, DPO, GRPO). This lab is the world in between.

Sources

  • Hugging Face, Building RL Environments with OpenEnv, Modules 1–5 (commit 57ea985, March 2026): the RL loop, the 3-method interface, typed models, deployment, the Module 4 word game, GRPO with TRL
  • meta-pytorch/OpenEnv, main (26 Sep 2026): src/openenv/core/env_server (create_fastapi_app, stateless HTTP routes, max_concurrent_envs), core/rubrics, rfcs/004-rubrics.md, rfcs/005-agentic-harnesses.md; PyPI openenv-core 0.3.0, against which the word game was run
  • TRL documentation, OpenEnv Integration for Training LLMs with Environments: environment_factory, "Tips for reward functions" (binary rewards, check the final state, test against a random baseline), max_completion_length across turns, server concurrency
  • The local lab behind this page (_study/openenv-course/lab): Frequency wins on neural and tensor only; the Farmer's 4.92 / 0.087 / 0 under the three rules; 1,000 games of Frequency pooled to 20.0%; the A/A false-alarm rate by 40,000-trial simulation
  • Pollen Robotics, microduck_rl (mjlab): 4,096 parallel environments, PPO with GAE λ = 0.95, actor and critic observations
  • "You're Not AI-Fluent Until You Understand Evals" (LangChain, 2026): verifier, environment, eval suite, harness engineering, trace

The sibling labs: Policy Pond (how a robot learns) and Finishing School (how a language model gets finished).