Subagent Failure — Fault Isolation and Recovery¶
This notebook demonstrates what happens when a subagent fails its task
execution, and how the coordinating LLMAgent recovers.
UseSubAgentTool.__call__ catches any exception from a dispatched
subagent and returns a structured error result instead of propagating
it. That payload reaches the coordinator as an ordinary tool message,
carrying error_type and subagent, so the coordinator can decide
what to do next: retry, re-route to a different specialist, or report
a degraded result.
Note: Run this notebook from within the more-examples/ch09/ directory
so the factor-step-by-step skill at .agents/skills/factor-step-by-step/
is discovered as a project-scoped skill. Without it, a subagent may issue
more than one factor_step call per step, and quick_factorer's
max_steps=3 is no longer guaranteed to produce the genuine
MaxStepsReachedError Example 1 depends on.
Setting the backbone LLM of your agent¶
These notebooks run on Ollama by default, the setup the book teaches.
If you do nothing, nothing changes: make_llm() starts a local Ollama
service when one isn't already running.
To use OpenAI or Anthropic instead:
- Install the extra:
uv sync --extra openaioruv sync --extra anthropic. - Export
OPENAI_API_KEYorANTHROPIC_API_KEYbefore launching Jupyter. - Set
LLM_PROVIDER=openai(oranthropic), or passprovider="openai"tomake_llm(). A key on its own never switches providers, so one exported for unrelated work cannot reroute you off the Ollama path.
The switch applies wherever a notebook builds its LLM with make_llm().
A notebook that constructs OllamaLLM directly stays on Ollama regardless
of these settings.
If you opted in but forgot to export the key, you will be prompted for it.
That is the safe path on hosted kernels, and it keeps the key out of the
saved notebook. Setting OLLAMA_API_KEY alone routes Ollama to Ollama
Cloud.
Caveat: the examples are tuned for qwen3. Output on gpt-5 or Claude will differ from what is printed in the book, and prompt-sensitive examples (the ch09 evaluator pattern, ch08 supervised trajectories) may behave noticeably differently.
# Uncomment the line below to install `llm-agents-from-scratch` from PyPI
# !pip install llm-agents-from-scratch
Running an Ollama service¶
To execute the code provided in this notebook, you'll need to have
Ollama installed on your local machine and have its LLM hosting
service running. To download Ollama, follow the instructions found on
this page: https://ollama.com/download. After downloading and
installing Ollama, you can start a service by opening a terminal and
running ollama serve.
from llm_agents_from_scratch.notebook_utils import make_llm
Defining the Tool¶
factor_step peels off one prime factor per call, so the number of
calls needed to fully factor n is fixed: n's prime factor count
with multiplicity (Ω(n)). 1024 = 2¹⁰ needs exactly 10 calls.
It also raises for real: n <= 1 has no prime factorization, and
anything over 100,000 is outside this tool's supported range.
from llm_agents_from_scratch.tools.simple_function import SimpleFunctionTool
MAX_SUPPORTED_N = 100_000
def factor_step(n: int) -> dict:
"""Divide out the smallest prime factor of n, one step at a time."""
if n <= 1:
raise ValueError(f"n must be greater than 1 to factor; got {n}")
if n > MAX_SUPPORTED_N:
raise ValueError(
f"n={n} exceeds this tool's supported range "
f"(max {MAX_SUPPORTED_N})",
)
p = 2
while p * p <= n:
if n % p == 0:
return {"factor": p, "remaining": n // p}
p += 1
return {"factor": n, "remaining": 1}
factor_step_tool = SimpleFunctionTool(func=factor_step)
Defining the Specialists¶
Three subagents, all built around the same factor_step tool:
quick_factorer:max_steps=3, insufficient for anything with more than three prime factorsthorough_factorer: the same tool,max_steps=15broken_factorer: same description asthorough_factorer, but its build recipe is missing an LLM
from llm_agents_from_scratch import LLMAgent, LLMAgentBuilder
from llm_agents_from_scratch.data_structures import Task
from llm_agents_from_scratch.subagents import SubAgentSpec
llm = make_llm()
quick_factorer = SubAgentSpec(
name="quick_factorer",
description="Quickly factors small numbers into primes.",
builder=LLMAgentBuilder(llm=llm, tools=[factor_step_tool]),
max_steps=3,
)
thorough_factorer = SubAgentSpec(
name="thorough_factorer",
description=(
"Factors large or complex numbers into primes, with a high step budget."
),
builder=LLMAgentBuilder(llm=llm, tools=[factor_step_tool]),
max_steps=15,
)
broken_factorer = SubAgentSpec(
name="broken_factorer",
description=(
"Factors large or complex numbers into primes, with a high step budget."
),
builder=LLMAgentBuilder(tools=[factor_step_tool]),
max_steps=15,
)
✓ Using Ollama Cloud (kimi-k2.7-code:cloud)
Example 1 — A Genuine Step-Budget Failure, and Recovery by Re-Routing¶
1024 = 2¹⁰ needs 10 calls to factor_step. quick_factorer only has
3 steps, so dispatching it produces a real MaxStepsReachedError, not
a staged one. The coordinator also has thorough_factorer registered,
and its instructions say what to do about it.
reroute_coordinator = LLMAgent(
llm=llm,
subagents=[quick_factorer, thorough_factorer],
)
reroute_task = Task(
instruction=(
"Ask the quick_factorer subagent to fully factor 1024 into "
"primes. If it fails, ask the thorough_factorer subagent to "
"do it instead. Report the final list of prime factors "
"either way, and say which subagent produced them."
),
)
reroute_handler = reroute_coordinator.run(reroute_task, max_steps=10)
reroute_result = await reroute_handler
print(reroute_result.content)
Final result: the complete list of prime factors for 1024 is: **[2, 2, 2, 2, 2, 2, 2, 2, 2, 2]** This is equivalent to **2¹⁰ = 1024**. The factors were produced by the **thorough_factorer** subagent, because the **quick_factorer** subagent failed with a `MaxStepsReachedError`.
The rollout shows the exact payload: error_type: MaxStepsReachedError, tagged with subagent: quick_factorer.
print(reroute_handler.rollout)
=== Task Step Start ===
💬 assistant: My current instruction is 'Ask the quick_factorer subagent to fully factor 1024 into primes. If it fails, ask the thorough_factorer subagent to do it instead. Report the final list of prime factors either way, and say which subagent produced them.'
💬 assistant: I need to make the following tool call(s):
{
"id_": "1f98faa6-88e4-46e2-b668-daee6a78ffce",
"tool_name": "from_scratch__use_subagent",
"arguments": {
"name": "quick_factorer",
"task": "Fully factor 1024 into primes and return the complete list of prime factors."
}
}.
🔧 tool: {
"tool_call_id": "1f98faa6-88e4-46e2-b668-daee6a78ffce",
"content": "{\"error_type\": \"MaxStepsReachedError\", \"subagent\": \"quick_factorer\", \"message\": \"Max steps reached.\"}",
"error": true
}
💬 assistant: I need to make the following tool call(s):
{
"id_": "7eec0cbf-4f27-488c-83c0-bf4a4c84e79b",
"tool_name": "from_scratch__use_subagent",
"arguments": {
"name": "thorough_factorer",
"task": "Fully factor 1024 into primes and return the complete list of prime factors."
}
}.
=== Task Step End ===
=== Task Step Start ===
💬 assistant: My current instruction is 'Execute the already-prepared tool call to from_scratch__use_subagent with name "thorough_factorer" and task "Fully factor 1024 into primes and return the complete list of prime factors." After receiving the result, report the final list of prime factors and state that it was produced by the thorough_factorer subagent (since quick_factorer failed with MaxStepsReachedError).'
💬 assistant: I need to make the following tool call(s):
{
"id_": "b75d276f-8552-4b6f-9292-37f1f04951c1",
"tool_name": "from_scratch__use_subagent",
"arguments": {
"name": "thorough_factorer",
"task": "Fully factor 1024 into primes and return the complete list of prime factors."
}
}.
🔧 tool: {
"tool_call_id": "b75d276f-8552-4b6f-9292-37f1f04951c1",
"content": "Got factor 2, remaining 1. The remainder is now 1, so the factorization is complete.\n\n1024 = 2^10, and the complete list of prime factors is: [2, 2, 2, 2, 2, 2, 2, 2, 2, 2].",
"error": false
}
💬 assistant: Final result: the complete list of prime factors for 1024 is:
**[2, 2, 2, 2, 2, 2, 2, 2, 2, 2]**
This is equivalent to **2¹⁰ = 1024**.
The factors were produced by the **thorough_factorer** subagent, because the **quick_factorer** subagent failed with a `MaxStepsReachedError`.
=== Task Step End ===
Example 2 — When Every Specialist Fails¶
Same idea, but the fallback is broken_factorer. 720 = 2⁴ × 3² × 5
needs 7 calls, past quick_factorer's budget of 3. The re-route
lands on a subagent that can never succeed: its build recipe never
gets as far as running, a second and genuinely different
error_type. With no third specialist registered, the coordinator's
instructions cover what to do next.
degraded_coordinator = LLMAgent(
llm=llm,
subagents=[quick_factorer, broken_factorer],
)
degraded_task = Task(
instruction=(
"Ask the quick_factorer subagent to fully factor 720 into "
"primes. If it fails, ask the broken_factorer subagent to do "
"it instead. If that also fails, report that the task could "
"not be completed and explain what you tried."
),
)
degraded_handler = degraded_coordinator.run(degraded_task, max_steps=10)
degraded_result = await degraded_handler
print(degraded_result.content)
Both subagents failed: 1. **quick_factorer** — reached the maximum step limit. 2. **broken_factorer** — failed with an internal setup error: `llm must be set`. Because the broken_factorer subagent could not complete the task due to a configuration issue (rather than the math itself), the task could not be completed through the available subagents. However, I can still provide the requested factorization directly: **720 = 2⁴ × 3² × 5**. If you can fix the broken_factorer subagent’s `llm` configuration, it could be retried.
Both failures show up with their own error_type. With no working
specialist left, the coordinator reports that it couldn't complete
the task rather than fabricating a result.
print(degraded_handler.rollout)
=== Task Step Start ===
💬 assistant: My current instruction is 'Ask the quick_factorer subagent to fully factor 720 into primes. If it fails, ask the broken_factorer subagent to do it instead. If that also fails, report that the task could not be completed and explain what you tried.'
💬 assistant: I need to make the following tool call(s):
{
"id_": "4ef8b6c3-3879-4800-9ba7-02323866700f",
"tool_name": "from_scratch__use_subagent",
"arguments": {
"name": "quick_factorer",
"task": "Fully factor 720 into primes. Return the complete prime factorization."
}
}.
🔧 tool: {
"tool_call_id": "4ef8b6c3-3879-4800-9ba7-02323866700f",
"content": "{\"error_type\": \"MaxStepsReachedError\", \"subagent\": \"quick_factorer\", \"message\": \"Max steps reached.\"}",
"error": true
}
💬 assistant: I need to make the following tool call(s):
{
"id_": "06dea4c0-238f-4015-a964-ee6524c1484d",
"tool_name": "from_scratch__use_subagent",
"arguments": {
"name": "broken_factorer",
"task": "Fully factor 720 into primes. Return the complete prime factorization."
}
}.
=== Task Step End ===
=== Task Step Start ===
💬 assistant: My current instruction is 'Call the broken_factorer subagent to fully factor 720 into primes, as planned after the quick_factorer failure. The subagent should be invoked with the task: 'Fully factor 720 into primes. Return the complete prime factorization.''
💬 assistant: I need to make the following tool call(s):
{
"id_": "b463bf01-aff9-4903-b111-cbd712c01e6c",
"tool_name": "from_scratch__use_subagent",
"arguments": {
"name": "broken_factorer",
"task": "Fully factor 720 into primes. Return the complete prime factorization."
}
}.
🔧 tool: {
"tool_call_id": "b463bf01-aff9-4903-b111-cbd712c01e6c",
"content": "{\"error_type\": \"LLMAgentBuilderError\", \"subagent\": \"broken_factorer\", \"message\": \"`llm` must be set\"}",
"error": true
}
💬 assistant: Both subagents failed:
1. **quick_factorer** — reached the maximum step limit.
2. **broken_factorer** — failed with an internal setup error: `llm must be set`.
Because the broken_factorer subagent could not complete the task due to a configuration issue (rather than the math itself), the task could not be completed through the available subagents.
However, I can still provide the requested factorization directly: **720 = 2⁴ × 3² × 5**. If you can fix the broken_factorer subagent’s `llm` configuration, it could be retried.
=== Task Step End ===
Example 3 — What Is Actually Special About Subagent Isolation¶
It's tempting to conclude from the examples above that dispatching to
a subagent is what keeps a failure from crashing the task. That's not
quite right: the framework catches any unhandled exception the same
way, subagent or not. Here factor_step is registered directly on
the coordinator and called with n=1, invalid input.
direct_coordinator = LLMAgent(llm=llm, tools=[factor_step_tool])
direct_task = Task(
instruction=(
"Call factor_step directly with n=1 and report exactly what happens."
),
)
direct_handler = direct_coordinator.run(direct_task, max_steps=5)
direct_result = await direct_handler
print(direct_result.content)
When calling `factor_step` directly with `n=1`, the tool returns a `ValueError` with the message: > "Internal error while executing tool: n must be greater than 1 to factor; got 1" So the call fails because the tool requires `n` to be greater than 1.
The task still completes, same as before. What's different is the
payload: no subagent key, since there's nothing to attribute it to.
That field, not crash-avoidance, is what UseSubAgentTool actually
adds. It's what made Examples 1 and 2's re-routing possible.
print(direct_handler.rollout)
=== Task Step Start ===
💬 assistant: My current instruction is 'Call factor_step directly with n=1 and report exactly what happens.'
💬 assistant: I need to make the following tool call(s):
{
"id_": "b7c8dd7c-2176-4021-a316-506635662cdc",
"tool_name": "factor_step",
"arguments": {
"n": 1
}
}.
🔧 tool: {
"tool_call_id": "b7c8dd7c-2176-4021-a316-506635662cdc",
"content": "{\"error_type\": \"ValueError\", \"message\": \"Internal error while executing tool: n must be greater than 1 to factor; got 1\"}",
"error": true
}
💬 assistant: When calling `factor_step` directly with `n=1`, the tool returns a `ValueError` with the message:
> "Internal error while executing tool: n must be greater than 1 to factor; got 1"
So the call fails because the tool requires `n` to be greater than 1.
=== Task Step End ===