Evaluator Pattern — Build, Review, and Revise¶
This notebook demonstrates the evaluator MAS pattern: a build_agent
produces a result, a review_agent checks it against a checklist, and
the coordinator loops build to review to revise until the checklist
passes.
Setting the backbone LLM of your agent¶
These notebooks run on Ollama by default, the setup the book teaches.
If you do nothing, nothing changes: make_llm() starts a local Ollama
service when one isn't already running.
To use OpenAI or Anthropic instead:
- Install the extra:
uv sync --extra openaioruv sync --extra anthropic. - Export
OPENAI_API_KEYorANTHROPIC_API_KEYbefore launching Jupyter. - Set
LLM_PROVIDER=openai(oranthropic), or passprovider="openai"tomake_llm(). A key on its own never switches providers, so one exported for unrelated work cannot reroute you off the Ollama path.
The switch applies wherever a notebook builds its LLM with make_llm().
A notebook that constructs OllamaLLM directly stays on Ollama regardless
of these settings.
If you opted in but forgot to export the key, you will be prompted for it.
That is the safe path on hosted kernels, and it keeps the key out of the
saved notebook. Setting OLLAMA_API_KEY alone routes Ollama to Ollama
Cloud.
Caveat: the examples are tuned for qwen3. Output on gpt-5 or Claude will differ from what is printed in the book, and prompt-sensitive examples (the ch09 evaluator pattern, ch08 supervised trajectories) may behave noticeably differently.
# Uncomment the line below to install `llm-agents-from-scratch` from PyPI
# !pip install llm-agents-from-scratch
Running an Ollama service¶
To execute the code provided in this notebook, you'll need to have
Ollama installed on your local machine and have its LLM hosting
service running. To download Ollama, follow the instructions found on
this page: https://ollama.com/download. After downloading and
installing Ollama, you can start a service by opening a terminal and
running ollama serve.
from llm_agents_from_scratch.notebook_utils import make_llm
Defining the Coding Standards¶
check_code is a house coding standard for a single function: it must
have a docstring, type hints on every parameter and the return value,
lines under 80 characters, and an exact required function name. This
mirrors what a real reviewer enforces: code that works isn't enough,
it also has to match the team's own conventions.
import ast
from llm_agents_from_scratch.tools.simple_function import SimpleFunctionTool
REQUIRED_NAME = "to_fahrenheit"
MAX_LINE_LEN = 79
def check_code(code: str) -> dict:
"""Check a function definition against the house coding standards."""
try:
tree = ast.parse(code)
except SyntaxError as e:
return {"compliant": False, "issues": [f"syntax error: {e}"]}
funcs = [n for n in ast.walk(tree) if isinstance(n, ast.FunctionDef)]
if not funcs:
return {"compliant": False, "issues": ["no function definition found"]}
fn = funcs[0]
issues = []
if fn.name != REQUIRED_NAME:
issues.append(
f"function must be named exactly {REQUIRED_NAME!r}, "
f"got {fn.name!r}",
)
if not ast.get_docstring(fn):
issues.append("function is missing a docstring")
if fn.returns is None:
issues.append("missing a return type annotation")
params = (
fn.args.posonlyargs
+ fn.args.args
+ fn.args.kwonlyargs
+ ([fn.args.vararg] if fn.args.vararg else [])
+ ([fn.args.kwarg] if fn.args.kwarg else [])
)
unannotated = [a.arg for a in params if a.annotation is None]
if unannotated:
issues.append(
f"missing type hints on parameter(s): {', '.join(unannotated)}",
)
long_lines = [
i + 1
for i, line in enumerate(code.splitlines())
if len(line) > MAX_LINE_LEN
]
if long_lines:
issues.append(f"line(s) exceed {MAX_LINE_LEN} characters: {long_lines}")
return {"compliant": not issues, "issues": issues}
check_code_tool = SimpleFunctionTool(func=check_code)
Defining the Specialists¶
build_agent only ever sees the brief below, never the standards
above. review_agent never sees them either, in the sense that it
doesn't recite them; it just calls check_code and reports whatever
the tool actually returns.
from llm_agents_from_scratch import LLMAgent, LLMAgentBuilder
from llm_agents_from_scratch.data_structures import Task
from llm_agents_from_scratch.subagents import SubAgentSpec
llm = make_llm()
build_agent = SubAgentSpec(
name="build_agent",
description="Writes Python functions to a given brief.",
builder=LLMAgentBuilder(llm=llm),
max_steps=5,
)
review_agent = SubAgentSpec(
name="review_agent",
description="Checks Python functions against the house coding standards.",
builder=LLMAgentBuilder(llm=llm, tools=[check_code_tool]),
max_steps=5,
)
✓ Using Ollama Cloud (kimi-k2.7-code:cloud)
Example — Converging on a Compliant Function¶
The brief below asks for a docstring, type hints, and a line-length
limit, but never says what the function must be named. check_code
still enforces an exact required name regardless, the kind of naming
convention a spec sheet states but an informal brief tends to leave
out.
brief = (
"Write a Python function that converts a Celsius temperature to "
"Fahrenheit. Include a docstring and type hints. Keep every line "
"under 80 characters. Return only the function code."
)
coordinator = LLMAgent(llm=llm, subagents=[build_agent, review_agent])
task = Task(
instruction=(
f"Ask build_agent to: {brief} "
"Then ask review_agent to check that exact code against the "
"house coding standards: quote the full code in review_agent's "
"task. If there are issues, ask build_agent to revise: quote "
"the current code and the specific issues (quote them exactly) "
"in build_agent's task. Then ask review_agent to check the "
"revised code the same way. Repeat until compliant, up to 3 "
"rounds. Call one subagent at a time and wait for its result "
"before calling the next one. Report the final code and "
"whether it's compliant."
),
)
handler = coordinator.run(task, max_steps=20)
result = await handler
print(result.content)
The review_agent confirmed the revised function is fully compliant. Here's the final result:
```python
def to_fahrenheit(celsius: float) -> float:
"""Convert a Celsius temperature to Fahrenheit.
Args:
celsius: The temperature in degrees Celsius.
Returns:
The equivalent temperature in degrees Fahrenheit.
"""
return (celsius * 9 / 5) + 32
```
**Status:** Compliant with house coding standards after one revision round.
handler.rollout shows why a revision was needed: review_agent's
tool result names the specific naming mismatch, straight from
check_code's own return value, not a guess, and the coordinator
relays that specific issue back to build_agent rather than just
asking it to try again.
print(handler.rollout)
=== Task Step Start ===
💬 assistant: My current instruction is 'Ask build_agent to: Write a Python function that converts a Celsius temperature to Fahrenheit. Include a docstring and type hints. Keep every line under 80 characters. Return only the function code. Then ask review_agent to check that exact code against the house coding standards: quote the full code in review_agent's task. If there are issues, ask build_agent to revise: quote the current code and the specific issues (quote them exactly) in build_agent's task. Then ask review_agent to check the revised code the same way. Repeat until compliant, up to 3 rounds. Call one subagent at a time and wait for its result before calling the next one. Report the final code and whether it's compliant.'
💬 assistant: I need to make the following tool call(s):
{
"id_": "0b447a5d-2764-4ac6-903d-9144e56d254c",
"tool_name": "from_scratch__use_subagent",
"arguments": {
"name": "build_agent",
"task": "Write a Python function that converts a Celsius temperature to Fahrenheit. Include a docstring and type hints. Keep every line under 80 characters. Return only the function code."
}
}.
🔧 tool: {
"tool_call_id": "0b447a5d-2764-4ac6-903d-9144e56d254c",
"content": "```python\ndef celsius_to_fahrenheit(celsius: float) -> float:\n \"\"\"Convert a Celsius temperature to Fahrenheit.\n\n Args:\n celsius: The temperature in degrees Celsius.\n\n Returns:\n The equivalent temperature in degrees Fahrenheit.\n \"\"\"\n return (celsius * 9 / 5) + 32\n```",
"error": false
}
💬 assistant: I need to make the following tool call(s):
{
"id_": "a9adc8d0-b994-4827-b19c-e7905012b400",
"tool_name": "from_scratch__use_subagent",
"arguments": {
"name": "review_agent",
"task": "Check the following Python code against the house coding standards and report any issues:\n\n```python\ndef celsius_to_fahrenheit(celsius: float) -> float:\n \"\"\"Convert a Celsius temperature to Fahrenheit.\n\n Args:\n celsius: The temperature in degrees Celsius.\n\n Returns:\n The equivalent temperature in degrees Fahrenheit.\n \"\"\"\n return (celsius * 9 / 5) + 32\n```"
}
}.
=== Task Step End ===
=== Task Step Start ===
💬 assistant: My current instruction is 'Execute the from_scratch__use_subagent tool call with name='review_agent' and the task quoted in the current_response to check the Celsius-to-Fahrenheit function against the house coding standards and report any issues. Wait for the result before proceeding.'
💬 assistant: I need to make the following tool call(s):
{
"id_": "5e3c4b64-1f75-4004-aa50-3d2aa6de180a",
"tool_name": "from_scratch__use_subagent",
"arguments": {
"name": "review_agent",
"task": "Check the following Python code against the house coding standards and report any issues:\n\n```python\ndef celsius_to_fahrenheit(celsius: float) -> float:\n \"\"\"Convert a Celsius temperature to Fahrenheit.\n\n Args:\n celsius: The temperature in degrees Celsius.\n\n Returns:\n The equivalent temperature in degrees Fahrenheit.\n \"\"\"\n return (celsius * 9 / 5) + 32\n```"
}
}.
🔧 tool: {
"tool_call_id": "5e3c4b64-1f75-4004-aa50-3d2aa6de180a",
"content": "The code does not pass the house coding standards.\n\n**Issue found:**\n- The function must be named exactly `to_fahrenheit`, but it is currently named `celsius_to_fahrenheit`.\n\nEverything else appears to be compliant. To fix it, simply rename the function to `to_fahrenheit`.",
"error": false
}
💬 assistant: I need to make the following tool call(s):
{
"id_": "e95a5116-9a71-436e-835a-fc45683d57e9",
"tool_name": "from_scratch__use_subagent",
"arguments": {
"name": "build_agent",
"task": "Revise the following Python function to address this specific issue. Return only the revised function code.\n\nCurrent code:\n```python\ndef celsius_to_fahrenheit(celsius: float) -> float:\n \"\"\"Convert a Celsius temperature to Fahrenheit.\n\n Args:\n celsius: The temperature in degrees Celsius.\n\n Returns:\n The equivalent temperature in degrees Fahrenheit.\n \"\"\"\n return (celsius * 9 / 5) + 32\n```\n\nIssue: \"The function must be named exactly `to_fahrenheit`, but it is currently named `celsius_to_fahrenheit`.\""
}
}.
=== Task Step End ===
=== Task Step Start ===
💬 assistant: My current instruction is 'Execute the from_scratch__use_subagent tool call with name='build_agent' and the task quoted in the current_response to revise the function by renaming it to 'to_fahrenheit'. Wait for the revised code before proceeding.'
💬 assistant: I need to make the following tool call(s):
{
"id_": "309f1c50-060b-420d-83a4-bc139ad9f0e3",
"tool_name": "from_scratch__use_subagent",
"arguments": {
"name": "build_agent",
"task": "Revise the following Python function to address this specific issue. Return only the revised function code.\n\nCurrent code:\n```python\ndef celsius_to_fahrenheit(celsius: float) -> float:\n \"\"\"Convert a Celsius temperature to Fahrenheit.\n\n Args:\n celsius: The temperature in degrees Celsius.\n\n Returns:\n The equivalent temperature in degrees Fahrenheit.\n \"\"\"\n return (celsius * 9 / 5) + 32\n```\n\nIssue: \"The function must be named exactly `to_fahrenheit`, but it is currently named `celsius_to_fahrenheit`.\""
}
}.
🔧 tool: {
"tool_call_id": "309f1c50-060b-420d-83a4-bc139ad9f0e3",
"content": "```python\ndef to_fahrenheit(celsius: float) -> float:\n \"\"\"Convert a Celsius temperature to Fahrenheit.\n\n Args:\n celsius: The temperature in degrees Celsius.\n\n Returns:\n The equivalent temperature in degrees Fahrenheit.\n \"\"\"\n return (celsius * 9 / 5) + 32\n```",
"error": false
}
💬 assistant: I need to make the following tool call(s):
{
"id_": "782d347e-4222-4ac3-a35b-5d6ae87abb0d",
"tool_name": "from_scratch__use_subagent",
"arguments": {
"name": "review_agent",
"task": "Check the following Python code against the house coding standards and report any issues:\n\n```python\ndef to_fahrenheit(celsius: float) -> float:\n \"\"\"Convert a Celsius temperature to Fahrenheit.\n\n Args:\n celsius: The temperature in degrees Celsius.\n\n Returns:\n The equivalent temperature in degrees Fahrenheit.\n \"\"\"\n return (celsius * 9 / 5) + 32\n```"
}
}.
=== Task Step End ===
=== Task Step Start ===
💬 assistant: My current instruction is 'Execute the from_scratch__use_subagent tool call with name='review_agent' and the task quoted in the current_response to check the revised to_fahrenheit function against the house coding standards. Wait for the result before proceeding.'
💬 assistant: I need to make the following tool call(s):
{
"id_": "4efebd0e-3a9e-4848-a8f2-7add06cc586b",
"tool_name": "from_scratch__use_subagent",
"arguments": {
"name": "review_agent",
"task": "Check the following Python code against the house coding standards and report any issues:\n\n```python\ndef to_fahrenheit(celsius: float) -> float:\n \"\"\"Convert a Celsius temperature to Fahrenheit.\n\n Args:\n celsius: The temperature in degrees Celsius.\n\n Returns:\n The equivalent temperature in degrees Fahrenheit.\n \"\"\"\n return (celsius * 9 / 5) + 32\n```"
}
}.
🔧 tool: {
"tool_call_id": "4efebd0e-3a9e-4848-a8f2-7add06cc586b",
"content": "The code is fully compliant with the house coding standards. No issues were found.",
"error": false
}
💬 assistant: The review_agent confirmed the revised function is fully compliant. Here's the final result:
```python
def to_fahrenheit(celsius: float) -> float:
"""Convert a Celsius temperature to Fahrenheit.
Args:
celsius: The temperature in degrees Celsius.
Returns:
The equivalent temperature in degrees Fahrenheit.
"""
return (celsius * 9 / 5) + 32
```
**Status:** Compliant with house coding standards after one revision round.
=== Task Step End ===
The 3-round cap is a safety net, not something this example needs.
review_agent never invents requirements, so once build_agent
finally sees the specific issue spelled out, satisfying it on the
next attempt is straightforward.