GAIA Agent — HF Agents Course Unit 4
A smolagents CodeAgent on free, open-weight models that answers GAIA Level-1 questions with web, Wikipedia, file, audio, image and video tools, and records every step it takes so each answer can be audited.
Graded runs, every step shown
Sixteen original GAIA-style tasks written for this project, each with a gold answer that can be checked. Because the tasks are ours, every question, answer and step of reasoning is public. Pick a run, then a task.
Step by step, what the agent did
Each step is one model call: the agent's reasoning, the Python it wrote, the tools that code called, and what came back, with timing and tokens per step. GAIA items appear as a step skeleton only.
Practice-set accuracy and ablations
Measured on the practice set with its own answer key, never on the course's GAIA questions. Repeated runs give pass@k and pass^k; ablations switch one thing off at a time.
What was submitted
One row per course question in the submitted run.
Where the effort went
Every model request, logged
The provider layer rate-limits, retries with backoff, falls back to the next free model, and logs each request. This is what that cost, across every run.
Where runs went wrong
Each trace is scanned for failure behaviours: timeouts, rate limits, crashed code, loops, empty answers. No answer key is involved, so a tag marks a risky behaviour, not a wrong answer. Every tag links to the trajectory that earned it.
Architecture
A single smolagents CodeAgent writes Python in a loop. Tools are plain functions it calls from that code. Every run leaves answers, traces and evals on disk, and this page is built from them.
Honesty note
- No answer lookups. The agent solves each question with its tools. Nothing is hard-coded, and prompts are never tuned to a known answer.
- GAIA guard. The web tools refuse the GAIA dataset and known answer-dump pages, so a search cannot return the key.
- No resharing. GAIA items appear on this page as ids, counts and tags only, under the contamination policy. The practice set is ours and is shown in full.
- Free, open-weight models only. Every model's weights are public on Hugging Face; served free, no paid API spend.
- Scores as returned. The official score is the course API's response, copied verbatim. The API grades only in aggregate, so per-question rows show what was submitted, not whether it was right.
- Small samples. Twenty questions is a small sample: a 95% interval around any score is roughly ±20 points. Ablation deltas smaller than that are noise until repeated.