Using Jev to decide the next step¶
Every step, TaskHandler.get_next_step() makes one structured_output()
call asking the backbone LLM to classify where things stand: next_step
or final_result. When it picks final_result, the LLM's own content
is discarded: the actual TaskResult.content comes from the previous
step's result, not from this call. So on the step that ends the task, a
full structured-output call is spent generating an answer that is thrown
away; only the kind label survives.
This notebook replaces that classification with a judgment call to
TypeSafe's Jev model instead. Jev's Noul primitive answers exactly this
shape of question: a calibrated yes/no judgment, returned as a
probability rather than a flat label. A JevTaskHandler asks Jev whether
the task is done; the LLM is only called to write the next instruction,
and only when Jev says the task isn't done yet. The step that used to
generate and discard an answer now makes no LLM call at all.
NOTE
LLMAgent.run() builds its own TaskHandler internally and has no
parameter for swapping in a subclass, and its driving loop
(_process_loop) is a closure with no extension point. This notebook
copies that loop's body into a manual while not task_handler.done()
loop driving a JevTaskHandler directly, the same pattern
run_supervised()'s callers already use for SupervisedTaskHandler.
Setting the backbone LLM of your agent¶
These notebooks run on Ollama by default, the setup the book teaches.
If you do nothing, nothing changes: make_llm() starts a local Ollama
service when one isn't already running.
To use OpenAI or Anthropic instead:
- Install the extra:
uv sync --extra openaioruv sync --extra anthropic. - Export
OPENAI_API_KEYorANTHROPIC_API_KEYbefore launching Jupyter. - Set
LLM_PROVIDER=openai(oranthropic), or passprovider="openai"tomake_llm(). A key on its own never switches providers, so one exported for unrelated work cannot reroute you off the Ollama path.
The switch applies wherever a notebook builds its LLM with make_llm().
A notebook that constructs OllamaLLM directly stays on Ollama regardless
of these settings.
If you opted in but forgot to export the key, you will be prompted for it.
That is the safe path on hosted kernels, and it keeps the key out of the
saved notebook. Setting OLLAMA_API_KEY alone routes Ollama to Ollama
Cloud.
Caveat: the examples are tuned for qwen3. Output on gpt-5 or Claude will differ from what is printed in the book, and prompt-sensitive examples (the ch09 evaluator pattern, ch08 supervised trajectories) may behave noticeably differently.
# Uncomment the line below to install `llm-agents-from-scratch` from PyPI
# !pip install llm-agents-from-scratch
This notebook uses TypeSafe SDK. To install it, run the command below.
!uv pip install typesafe-sdk -q
Requirements¶
- A TypeSafe API key stored in the environment variable
TYPESAFE_API_KEY, created at console.typesafe.ai
import getpass
import os
if not os.environ.get("TYPESAFE_API_KEY"):
os.environ["TYPESAFE_API_KEY"] = getpass.getpass("TypeSafe API Key: ")
import json
import logging
import urllib.request
from llm_agents_from_scratch import LLMAgent
from llm_agents_from_scratch.data_structures import Task
from llm_agents_from_scratch.logger import enable_console_logging
from llm_agents_from_scratch.notebook_utils import make_llm
from llm_agents_from_scratch.tools import SimpleFunctionTool
enable_console_logging(logging.INFO)
def get_next_evolution(species_name: str) -> str:
"""Looks up the next evolution stage for a Pok\u00e9mon species.
Returns the next stage's name, or "no further evolution" if the
species is already at the end of its evolution chain.
"""
species_name = species_name.lower().strip()
req = urllib.request.Request(
f"https://pokeapi.co/api/v2/pokemon-species/{species_name}",
headers={"User-Agent": "llm-agents-from-scratch/1.0"},
)
with urllib.request.urlopen(req, timeout=10) as resp:
species = json.loads(resp.read())
chain_req = urllib.request.Request(
species["evolution_chain"]["url"],
headers={"User-Agent": "llm-agents-from-scratch/1.0"},
)
with urllib.request.urlopen(chain_req, timeout=10) as resp:
chain = json.loads(resp.read())["chain"]
def _find(node):
if node["species"]["name"] == species_name:
return node
for child in node["evolves_to"]:
found = _find(child)
if found is not None:
return found
return None
node = _find(chain)
if node is None or not node["evolves_to"]:
return "no further evolution"
return node["evolves_to"][0]["species"]["name"]
llm = make_llm()
agent = LLMAgent(
llm=llm,
tools=[SimpleFunctionTool(func=get_next_evolution)],
)
TASK_INSTRUCTION = (
"Starting from charmander, use get_next_evolution (only, not your "
"prior knowledge) one step at a time to walk its evolution chain "
"until there is no further evolution, then report the complete "
"chain."
)
✓ Using Ollama Cloud (kimi-k2.7-code:cloud)
Baseline: the LLM decides¶
agent.run() drives the task with the stock TaskHandler: every step,
get_next_step() asks the LLM to classify the current response as
next_step or final_result in one structured_output() call.
baseline_task = Task(instruction=TASK_INSTRUCTION)
baseline_result = await agent.run(baseline_task)
print(baseline_result.content)
INFO (llm_agents_fs.LLMAgent) : 🚀 Starting task: Starting from charmander, use get_next_evolution (only, not your prior knowledge) one step at a time to walk its evolution chain unti...[TRUNCATED]
INFO (llm_agents_fs.TaskHandler) : ⚙️ Processing Step: Starting from charmander, use get_next_evolution (only, not your prior knowledge) one step at a time to walk its evolution chain u...[TRUNCATED]
INFO (llm_agents_fs.TaskHandler) : 🛠️ Executing Tool Call: get_next_evolution
INFO (llm_agents_fs.TaskHandler) : ✅ Successful Tool Call: charmeleon
INFO (llm_agents_fs.TaskHandler) : ✅ Step Result: I need to make the following tool-calls:
{
"id_": "3432d5da-5747-49c0-b3d5-b8fbceae0f25",
"tool_name": "get_next_evolution",
...[TRUNCATED]
INFO (llm_agents_fs.TaskHandler) : 🧠 New Step: Execute the get_next_evolution tool call for charmeleon and continue walking the evolution chain one step at a time until no further evolu...[TRUNCATED]
INFO (llm_agents_fs.TaskHandler) : ⚙️ Processing Step: Execute the get_next_evolution tool call for charmeleon and continue walking the evolution chain one step at a time until no furth...[TRUNCATED]
INFO (llm_agents_fs.TaskHandler) : 🛠️ Executing Tool Call: get_next_evolution
INFO (llm_agents_fs.TaskHandler) : ✅ Successful Tool Call: charizard
INFO (llm_agents_fs.TaskHandler) : ✅ Step Result: I need to make the following tool-calls:
{
"id_": "debc115d-e7d3-4983-af54-89af6109286f",
"tool_name": "get_next_evolution",
...[TRUNCATED]
INFO (llm_agents_fs.TaskHandler) : 🧠 New Step: Execute the pending get_next_evolution tool call for 'charizard'. If it returns another evolution, continue calling get_next_evolution one...[TRUNCATED]
INFO (llm_agents_fs.TaskHandler) : ⚙️ Processing Step: Execute the pending get_next_evolution tool call for 'charizard'. If it returns another evolution, continue calling get_next_evolu...[TRUNCATED]
INFO (llm_agents_fs.TaskHandler) : 🛠️ Executing Tool Call: get_next_evolution
INFO (llm_agents_fs.TaskHandler) : ✅ Successful Tool Call: no further evolution
INFO (llm_agents_fs.TaskHandler) : ✅ Step Result: The tool returned "no further evolution" for charizard, so the chain is complete.
Starting from charmander, the full evolution chain i...[TRUNCATED]
INFO (llm_agents_fs.TaskHandler) : No new step required.
INFO (llm_agents_fs.LLMAgent) : 🏁 Task completed: The tool returned "no further evolution" for charizard, so the chain is complete.
Starting from charmander, the full evolution chai...[TRUNCATED]
The tool returned "no further evolution" for charizard, so the chain is complete.
Starting from charmander, the full evolution chain is:
1. charmander → charmeleon
2. charmeleon → charizard
3. charizard → no further evolution
Complete chain: **charmander → charmeleon → charizard**
Splitting judgment from generation: JevTaskHandler¶
JevTaskHandler overrides get_next_step(). Instead of one LLM call
that both classifies and (possibly wastefully) generates, it asks Jev's
Noul primitive a single yes/no question (is the task done) over a
state built from the same rollout, current response, and instruction
the stock prompt uses. Jev returns a probability, compared against
done_threshold:
- Done: no LLM call at all. The previous step's content becomes the
TaskResultdirectly, exactly like the stock handler, just without the discarded generation. - Not done: the LLM is called once with a continuation-only prompt
and a schema that only has
content-- nokindfield at all, so it can't disagree with Jev by classifyingfinal_resultand returning an empty instruction, the way reusing the stockNextStepDecisionschema could.
from typing import Any
from pydantic import BaseModel, Field
from typesafe_sdk import AsyncTypeSafeClient, Noul, NoulCriteria
from llm_agents_from_scratch.data_structures import (
RejectedTaskResult,
TaskResult,
TaskStep,
TaskStepResult,
)
class NextStepContent(BaseModel):
"""Structured output used by JevTaskHandler.get_next_step()."""
content: str = Field(
description="The next tool call or action to take.",
)
class JevTaskHandler(LLMAgent.TaskHandler):
"""A TaskHandler whose stop/continue judgment comes from Jev.
Attributes:
typesafe_client (AsyncTypeSafeClient): Client used to reach Jev.
done_threshold (float): Minimum Noul probability to treat the
task as complete.
"""
def __init__(
self,
*args: Any,
typesafe_client: AsyncTypeSafeClient,
done_threshold: float = 0.8,
**kwargs: Any,
) -> None:
"""Initialize a JevTaskHandler.
Args:
typesafe_client (AsyncTypeSafeClient): Client used to reach
Jev.
done_threshold (float): Minimum Noul probability to treat
the task as complete. Defaults to 0.8.
*args: Forwarded to TaskHandler.
**kwargs: Forwarded to TaskHandler.
"""
super().__init__(*args, **kwargs)
self.typesafe_client = typesafe_client
self.done_threshold = done_threshold
async def get_next_step(
self,
previous_step_result: TaskStepResult | RejectedTaskResult | None,
) -> TaskStep | TaskResult:
"""Ask Jev whether the task is done; the LLM only writes steps."""
if not previous_step_result or isinstance(
previous_step_result,
RejectedTaskResult,
):
return await super().get_next_step(previous_step_result)
state = (
f"<user-instruction>\n{self.task.instruction}\n"
"</user-instruction>\n\n"
f"<thinking-process>\n{self.rollout}\n</thinking-process>\n\n"
f"<current-response>\n{previous_step_result.content}\n"
"</current-response>"
)
response = await self.typesafe_client.system_one(
state=state,
questions={
"is_done": Noul(
instructions=(
"Given the user's instruction, the assistant's "
"rollout so far, and its most recent response, "
"decide whether the task is fully complete."
),
criteria=NoulCriteria(
true=(
"The current response is a final answer; no "
"further tool calls or steps are needed."
),
false=(
"The current response describes a pending "
"action, a plan, or an intermediate result "
"-- more work is needed."
),
),
),
},
model="jev-latest",
)
confidence = response.answers["is_done"].noul
self.logger.info(
f"\U0001f52e Jev confidence task is done: {confidence:.2f}",
)
if confidence >= self.done_threshold:
return TaskResult(
task_id=self.task.id_,
content=previous_step_result.content,
)
prompt = (
"You are overseeing an assistant's progress in "
"accomplishing a user instruction. The task is not yet "
"complete; more steps are needed. Describe the next tool "
"call or action the assistant should take.\n\n"
f"<thinking-process>\n{self.rollout}\n</thinking-process>\n\n"
f"<current-response>\n{previous_step_result.content}\n"
"</current-response>\n\n"
f"<user-instruction>\n{self.task.instruction}\n"
"</user-instruction>"
)
next_content = await self.llm_agent.llm.structured_output(
prompt=prompt,
mdl=NextStepContent,
)
return TaskStep(
task_id=self.task.id_,
instruction=next_content.content,
)
Driving JevTaskHandler manually¶
This is _process_loop()'s body, unchanged, driving a JevTaskHandler
instead of the stock one.
We can run a task directly with JevTaskHandler, just by passing it an LLMAgent
and a Task to perform.
from llm_agents_from_scratch.errors import MaxStepsReachedError
async def run_with_jev(
agent: LLMAgent,
task: Task,
typesafe_client: AsyncTypeSafeClient,
done_threshold: float = 0.8,
max_steps: int | None = None,
) -> "LLMAgent.TaskHandler":
"""Drives a JevTaskHandler to completion, mirroring _process_loop()."""
task_handler = JevTaskHandler(
llm_agent=agent,
task=task,
typesafe_client=typesafe_client,
done_threshold=done_threshold,
)
await task_handler.load_memories()
step_result = None
while not task_handler.done():
try:
if task_handler.step_counter == max_steps:
raise MaxStepsReachedError("Max steps reached.")
next_step = await task_handler.get_next_step(step_result)
match next_step:
case TaskStep():
step_result = await task_handler.run_step(next_step)
case TaskResult():
await task_handler.record_memory(result=next_step)
task_handler.set_result(next_step)
except Exception as e:
await task_handler.record_memory(error=e)
task_handler.set_exception(e)
return task_handler
jev_task = Task(instruction=TASK_INSTRUCTION)
async with AsyncTypeSafeClient() as typesafe_client:
jev_task_handler = await run_with_jev(
agent,
jev_task,
typesafe_client,
max_steps=5,
)
jev_result = jev_task_handler.result()
print(jev_result.content)
INFO (llm_agents_fs.JevTaskHandler) : ⚙️ Processing Step: Starting from charmander, use get_next_evolution (only, not your prior knowledge) one step at a time to walk its evolution chain u...[TRUNCATED]
INFO (llm_agents_fs.JevTaskHandler) : 🛠️ Executing Tool Call: get_next_evolution
INFO (llm_agents_fs.JevTaskHandler) : ✅ Successful Tool Call: charmeleon
INFO (llm_agents_fs.JevTaskHandler) : ✅ Step Result: I need to make the following tool-calls:
{
"id_": "495a84f9-6de2-445b-9fcc-20409c4758b7",
"tool_name": "get_next_evolution",
...[TRUNCATED]
INFO (llm_agents_fs.JevTaskHandler) : 🔮 Jev confidence task is done: 0.02
INFO (llm_agents_fs.JevTaskHandler) : ⚙️ Processing Step: Call get_next_evolution with species_name "charmeleon".
INFO (llm_agents_fs.JevTaskHandler) : 🛠️ Executing Tool Call: get_next_evolution
INFO (llm_agents_fs.JevTaskHandler) : ✅ Successful Tool Call: charizard
INFO (llm_agents_fs.JevTaskHandler) : ✅ Step Result: I need to make the following tool-calls:
{
"id_": "4be06ec3-1a41-42a1-a399-5242f4f27217",
"tool_name": "get_next_evolution",
...[TRUNCATED]
INFO (llm_agents_fs.JevTaskHandler) : 🔮 Jev confidence task is done: 0.03
INFO (llm_agents_fs.JevTaskHandler) : ⚙️ Processing Step: Call get_next_evolution with species_name "charizard".
INFO (llm_agents_fs.JevTaskHandler) : 🛠️ Executing Tool Call: get_next_evolution
INFO (llm_agents_fs.JevTaskHandler) : ✅ Successful Tool Call: no further evolution
INFO (llm_agents_fs.JevTaskHandler) : ✅ Step Result: The tool returned "no further evolution", which means Charizard is the final stage. The complete evolution chain starting from Charmand...[TRUNCATED]
INFO (llm_agents_fs.JevTaskHandler) : 🔮 Jev confidence task is done: 0.98
The tool returned "no further evolution", which means Charizard is the final stage. The complete evolution chain starting from Charmander is:
**Charmander → Charmeleon → Charizard**
Comparing the two runs¶
Both runs reach the same answer, driven by the same backbone LLM for tool use and step generation. What changed is the next-step/stopping decision:
- The original
TaskHandlerasks the LLM to classify and generate on every step, discarding the generation whenever the classification isfinal_result. JevTaskHandlerasks Jev for a calibrated probability first, and only calls the LLM to generate when there's genuinely something left to write.done_thresholdturns that probability into a tunable policy: raise it to make the agent more conservative about declaring a task finished, lower it to stop sooner, a lever the binarykindlabel never offered.
It's not hard to think of where else Jev may be used in the processing loop. Essentially anywhere a structured decision needs to be made is a candidate application for Jev. Looking ahead to capabilities future chapters enable:
- Jev could judge whether a skill should be activated before the
single-turn chat with the backbone LLM even starts, rather than
leaving that decision to the LLM's own
use_skilltool call (Chapter 6). - Jev could judge whether a subagent should be dispatched for a given
step, rather than leaving that decision to the LLM's own
use_subagenttool call (Chapter 9).