Hugging Face Agents Course · Unit 4 · GAIA Level 1

GAIA Agent — HF Agents Course Unit 4

A smolagents CodeAgent on free, open-weight models that answers GAIA Level-1 questions with web, Wikipedia, file, audio, image and video tools, and records every step it takes so each answer can be audited.

Practice set · our own tasks

Graded runs, every step shown

Sixteen original GAIA-style tasks written for this project, each with a gold answer that can be checked. Because the tasks are ours, every question, answer and step of reasoning is public. Pick a run, then a task.

Trajectory viewer

Step by step, what the agent did

Each step is one model call: the agent's reasoning, the Python it wrote, the tools that code called, and what came back, with timing and tokens per step. GAIA items appear as a step skeleton only.

Evals

Practice-set accuracy and ablations

Measured on the practice set with its own answer key, never on the course's GAIA questions. Repeated runs give pass@k and pass^k; ablations switch one thing off at a time.

GAIA run · the course's 20 questions

What was submitted

One row per course question in the submitted run.

Analytics · GAIA run

Where the effort went

Free-tier usage

Every model request, logged

The provider layer rate-limits, retries with backoff, falls back to the next free model, and logs each request. This is what that cost, across every run.

Failure taxonomy · GAIA run

Where runs went wrong

Each trace is scanned for failure behaviours: timeouts, rate limits, crashed code, loops, empty answers. No answer key is involved, so a tag marks a risky behaviour, not a wrong answer. Every tag links to the trajectory that earned it.

How it works

Architecture

A single smolagents CodeAgent writes Python in a loop. Tools are plain functions it calls from that code. Every run leaves answers, traces and evals on disk, and this page is built from them.

Honesty note

  • No answer lookups. The agent solves each question with its tools. Nothing is hard-coded, and prompts are never tuned to a known answer.
  • GAIA guard. The web tools refuse the GAIA dataset and known answer-dump pages, so a search cannot return the key.
  • No resharing. GAIA items appear on this page as ids, counts and tags only, under the contamination policy. The practice set is ours and is shown in full.
  • Free, open-weight models only. Every model's weights are public on Hugging Face; served free, no paid API spend.
  • Scores as returned. The official score is the course API's response, copied verbatim. The API grades only in aggregate, so per-question rows show what was submitted, not whether it was right.
  • Small samples. Twenty questions is a small sample: a 95% interval around any score is roughly ±20 points. Ablation deltas smaller than that are noise until repeated.