Policy Pond taught the learner and Finishing School taught the LLM learner. Neither taught the world they practice in. This is that world: a sealed cabinet with three controls, a ticket counter inside, and a strict rule about what the player may see. About 75 minutes of play, with a real environment running in your browser.
An environment is the world an agent practices in. For an LLM, OpenEnv packages that world as a sealed box running in its own container, with exactly three controls. Every environment has the same three, whether it plays Hangman, Wordle, a code repo or a fake CRM. Every term with a wavy red underline is tappable: hover or tap it for the plain meaning and its arcade equivalent.
Covers OpenEnv course Module 1 · the RL loop, the 3-method interface · Module 2 · typed models
This is the course's own Module 4 game, running here logic for logic. A hidden six-letter word; one letter per move; ten wrong guesses and you lose. Drop a coin, then press letters. The tape below prints exactly what each call returned.
/reset and /step are stateless: each call builds a fresh environment, answers, and throws it away. Only a WebSocket session keeps one game alive between calls, which is why the typed client uses /ws. Verified on the real server: two one-off steps of the same wrong letter both reported 9 attempts left.app.py raises max_concurrent_envs, so a trainer that opens 64 connections to a default server fails. Set it to at least the generation batch size.The most important design decision in any environment is made in its types file: which fields go on the observation (the front glass, what the player sees), which go in the state (the back panel: bookkeeping the client can still fetch), and which never leave the server. In RL terms the state is the world's full true situation and the observation is the player's partial view of it. The course's word game puts the secret word in state. Sort the fields yourself, then let a cheater play against your layout.
Covers Module 4 · Step 1, models.py · Module 1 · observation vs state · the study pack's M4 gotcha 5
Tap a field, then tap a bin. (On a computer you can also drag.) The cabinet needs the word to run the game; the question is where it may sit.
env.state() returns to the clienttarget_word in the State model. Can a player read it?env.state() is one of the three client methods, and the server answers it with the whole State model. Verified against the real server: state() came back with target_word='tensor' mid-game. A training loop that ever puts state() in the prompt, or an agent with a tool that calls it, learns to read the answer instead of playing.target_word in the client's _parse_state. Is the leak closed?/schema in twenty lines. A secret has to stay on the server, as a private attribute the game logic reads and nothing serializes. Fix the server's reset(), not the client.An optimizer doesn't play politely. It sends whatever pays. So before anything trains against a cabinet, send it the moves a good player never would and read what comes back. The course's game has four such holes, all of them found by running it. Each switch below closes one.
Covers Module 4 · Step 2, environment.py · TRL's "raise on invalid actions" rule · the study pack's M4 gotchas 4, 8 and 9
reset() scores 100% for nothing.A trainer doesn't know what you meant. It raises the average tickets. So for any reward rule the only question is: which player does it rank first? If that isn't the player you want, training will find the one you don't. Three rules, three players. The Farmer never wins a game; it finds one correct letter and presses it forever.
Covers Module 5 · reward functions, shaped vs binary · TRL's OpenEnv docs, "Tips for reward functions" · Finishing School step 6 (the group is the critic)
GRPO scores a group of games on the same prompt against their own average. The course trains with groups of 2. If the two games score the same, the advantage is zero and that group teaches nothing. Same player, same word, two games: how often do they score differently under the rule above?
An eval is the cabinet with a scoreboard: a fixed set of games, a verifier inside, and a pass rate for a human to read. Module 2's "policy competition" was one, without the two things that make a score believable: a floor and a ceiling, and error bars. This scoreboard has both, plus an A/A test: two identical Random players who must post the same score, or the harness itself is broken.
Covers Module 2 · policies and the competition · the LangChain evals vocabulary (verifier, environment, harness) · TRL's "if a capable model can't beat random…"
The course argues that in-process environments (Gym-style) are "a disaster for production." That's true for LLM training, where every step waits ~200 ms for the model to think and the game itself may be a code sandbox you don't want inside your trainer. It's the opposite for Microduck, whose 4,096 ducklings advance together as one tensor on one GPU; a network hop per step would cost more than the whole step. Move the sliders and watch which term owns the wall clock.
Covers Module 1 · why Gym falls short · the course's scaling appendix (WebSocket vs HTTP) · Policy Pond step 6 (Microduck's real config)
Three environments came back from training with suspicious scores. For each, read the symptom and the evidence, name the cause, then pick the fix. The tempting wrong fixes are there too, and each fails for a stated reason.
Covers everything above · the study pack's evals bridge §4 checklist
Eight questions, each answered by setting up a widget in steps 0–5 the way the question says and reading off the result. Graded instantly, with rounding tolerance.
Fix all three cabinets on the review board and score 6 or better here, and the arcade posts your high score.
Exact: the Hangman environment is the OpenEnv course's Module 4 game, logic for logic, including its bugs, and its behaviour on the real server (openenv-core 0.3.0) was checked call by call before this page was written. Frequency's and the Farmer's averages are exact expectations: one game per word over the course's ten words, drawn uniformly. The standard errors and the A/A test are the plain binomial formulas.
Simulated: Random's and the Cheater's scores, and the "two games differ" rates, come from a seeded pseudo-random generator, so they repeat exactly across reloads but are samples, not truths (Random's true win rate is about 2.3%). Illustrative: step 5's times are order-of-magnitude defaults on sliders; the review board's three cases are composites, not incidents.
The environment is real; the numbers about it were measured; the timing model is a sketch.
rfcs/ folder explains rubrics (composable rewards) and agentic harnesses.environment_factory, the reward-function tips this page quotes, and the concurrency note.main (26 Sep 2026): src/openenv/core/env_server (create_fastapi_app, stateless HTTP routes, max_concurrent_envs), core/rubrics, rfcs/004-rubrics.md, rfcs/005-agentic-harnesses.md; PyPI openenv-core 0.3.0, against which the word game was runenvironment_factory, "Tips for reward functions" (binary rewards, check the final state, test against a random baseline), max_completion_length across turns, server concurrency_study/openenv-course/lab): Frequency wins on neural and tensor only; the Farmer's 4.92 / 0.087 / 0 under the three rules; 1,000 games of Frequency pooled to 20.0%; the A/A false-alarm rate by 40,000-trial simulationmicroduck_rl (mjlab): 4,096 parallel environments, PPO with GAE λ = 0.95, actor and critic observationsThe sibling labs: Policy Pond (how a robot learns) and Finishing School (how a language model gets finished).