Cancelling a Task Execution — Stopping an In-Flight A2A Task¶
This notebook cancels a real, genuinely in-flight A2A task via
cancel(): dispatch a long Hailstone sequence, cancel it mid-step, and
confirm both that the task settles TASK_STATE_CANCELED and that the
underlying work actually stopped, not just that its status changed.
NOTE
There is no client-side cancel in this framework yet: a2a/client/
has no cancel at all, so UseA2AAgentTool cannot cancel a peer. This
notebook drives the raw a2a-sdk client directly.
cancel()'s ordering matters here: it has to publish CANCELED before
anything that can suspend, or a live subscriber can miss it entirely
(a real, previously confirmed failure mode, not a hypothetical one;
see LLMAgentA2AExecutor.cancel()'s own docstring). One open, known
gap this notebook doesn't exercise: if the agent completes normally at
the same instant a cancel arrives, both CANCELED and COMPLETED can
end up published for the same task, a separate issue already tracked
elsewhere, not something fixed here.
Setup Instructions¶
This notebook launches the extra/a2a-from-scratch-hailstone app from
a local checkout, so it needs the full repository, not just the
package. Launch it with:
uv run --with jupyter jupyter lab
from the project root. No PyPI release of llm-agents-from-scratch has
A2A support yet (still under "Unreleased" in CHANGELOG.md), so the
pip install cell below won't provide what this notebook needs even
if uncommented; it's left only for parity with the other notebooks.
Setting the backbone LLM of your agent¶
These notebooks run on Ollama by default, the setup the book teaches.
If you do nothing, nothing changes: make_llm() starts a local Ollama
service when one isn't already running.
To use OpenAI or Anthropic instead:
- Install the extra:
uv sync --extra openaioruv sync --extra anthropic. - Export
OPENAI_API_KEYorANTHROPIC_API_KEYbefore launching Jupyter. - Set
LLM_PROVIDER=openai(oranthropic), or passprovider="openai"tomake_llm(). A key on its own never switches providers, so one exported for unrelated work cannot reroute you off the Ollama path.
The switch applies wherever a notebook builds its LLM with make_llm().
A notebook that constructs OllamaLLM directly stays on Ollama regardless
of these settings.
If you opted in but forgot to export the key, you will be prompted for it.
That is the safe path on hosted kernels, and it keeps the key out of the
saved notebook. Setting OLLAMA_API_KEY alone routes Ollama to Ollama
Cloud.
Caveat: the examples are tuned for qwen3. Output on gpt-5 or Claude will differ from what is printed in the book, and prompt-sensitive examples (the ch09 evaluator pattern, ch08 supervised trajectories) may behave noticeably differently.
# No PyPI release has A2A support yet -- see Setup Instructions above.
# !pip install llm-agents-from-scratch
Running an Ollama service¶
To execute the code provided in this notebook, you'll need to have
Ollama installed on your local machine and have its LLM hosting
service running. To download Ollama, follow the instructions found on
this page: https://ollama.com/download. After downloading and
installing Ollama, you can start a service by opening a terminal and
running ollama serve.
import os
import subprocess
import time
import urllib.error
import urllib.request
from llm_agents_from_scratch.notebook_utils import ensure_ollama
use_cloud = "OLLAMA_API_KEY" in os.environ
ensure_ollama() if not use_cloud else print("\u2713 Using Ollama Cloud")
✓ Using Ollama Cloud
The Server¶
The same standalone app more-examples/ch10/streaming_llmagent_executor.ipynb
uses: extra/a2a-from-scratch-hailstone, a genuinely separate OS
process with its own log stream. 27 is the starting number: its
Hailstone sequence takes over 100 steps, leaving plenty of margin to
cancel well before it would ever finish on its own.
import contextlib
import signal
from pathlib import Path
from urllib.parse import urlparse
def find_repo_root(marker="extra"):
"""Walk up from cwd to the first ancestor containing `marker`.
Jupyter sets a kernel's cwd to the directory of the opened notebook
file under normal use, but that's not guaranteed for every launch
method (e.g. running via nbconvert from the project root), so this
doesn't assume a fixed number of parent hops.
"""
for candidate in (Path.cwd(), *Path.cwd().parents):
if (candidate / marker).is_dir():
return candidate
msg = f"Could not find a parent directory containing {marker!r}."
raise RuntimeError(msg)
def ensure_hailstone_server(host, log_path, timeout=15):
"""Start the from-scratch Hailstone A2A server app if not running.
Returns the Popen handle if this call started the server, or None
if it was already running (so Cleanup knows not to tear it down).
Runs in its own session (start_new_session=True) so Cleanup can
kill the whole process group -- terminating only the `uv run`
wrapper process can leave the actual uvicorn process orphaned and
still bound to the port.
"""
def _up():
try:
with urllib.request.urlopen(
f"{host}/.well-known/agent-card.json",
timeout=1,
):
pass
return True
except (urllib.error.URLError, ConnectionError, TimeoutError):
return False
if _up():
print(f"\u2713 A2A server already running at {host}")
return None
server_path = find_repo_root() / "extra" / "a2a-from-scratch-hailstone"
port = urlparse(host).port or 80
print(f"Starting A2A server at {host}...")
with open(log_path, "w") as log_file:
process = subprocess.Popen(
[
"uv",
"run",
"uvicorn",
"main:app",
"--host",
"0.0.0.0",
"--port",
str(port),
],
cwd=server_path,
stdout=log_file,
stderr=subprocess.STDOUT,
start_new_session=True,
)
deadline = time.time() + timeout
while time.time() < deadline:
if _up():
print(f"\u2713 A2A server up at {host}")
return process
time.sleep(0.5)
stop_hailstone_server(process)
raise RuntimeError(f"A2A server did not start within {timeout}s")
def stop_hailstone_server(process):
"""Kills the whole process group started by ensure_hailstone_server()."""
def _signal(sig):
with contextlib.suppress(ProcessLookupError):
os.killpg(os.getpgid(process.pid), sig)
_signal(signal.SIGTERM)
try:
process.wait(timeout=5)
except subprocess.TimeoutExpired:
_signal(signal.SIGKILL)
process.wait(timeout=5)
server_process = ensure_hailstone_server(
"http://127.0.0.1:9300",
"hailstone_server.log",
)
def stop_server():
"""Tears down the server this notebook started, if it started one.
Called from every later cell's except block so a failure partway
through -- discovery, cancellation, confirmation -- can't leave the
process running and the port occupied.
"""
if server_process is not None:
stop_hailstone_server(server_process)
Starting A2A server at http://127.0.0.1:9300... ✓ A2A server up at http://127.0.0.1:9300
A Fresh Client Per Role¶
Everything below opens its own client rather than sharing one: dispatch
gets a client, and cancellation gets a separate one later. This is
import asyncio plus one small helper, open_client(), that every
step below reuses.
import asyncio
from contextlib import asynccontextmanager
import httpx
from a2a.client import ClientConfig, create_client
from a2a.helpers import new_text_message
from a2a.types import (
CancelTaskRequest,
ListTasksRequest,
SendMessageRequest,
TaskState,
)
from a2a.types import (
Role as A2ARole,
)
from llm_agents_from_scratch.a2a import A2AAgentSpec
spec = await A2AAgentSpec.from_url("http://127.0.0.1:9300", timeout=120.0)
@asynccontextmanager
async def open_client():
"""A fresh SDK client on its own httpx connection every time.
Dispatch and cancel each open their own client below rather than
sharing one: a real canceller (an operator, a monitoring service)
is never the same connection that dispatched the task in the
first place.
"""
async with httpx.AsyncClient(timeout=120.0) as httpx_client:
client = await create_client(
agent=spec.agent_card,
client_config=ClientConfig(
streaming=False,
httpx_client=httpx_client,
),
)
try:
yield client
finally:
await client.close()
Step 1 — Dispatch the Task¶
send_message() doesn't return until the task is done, so it has to
run in the background (asyncio.ensure_future()) while the rest of
this notebook goes to cancel it. Nothing about this call returns a
task id directly, either: dispatch() only sees the final result,
whenever that ends up being.
task_text = (
"Compute the full hailstone sequence starting at 27, until it reaches 1."
)
message = new_text_message(text=task_text, role=A2ARole.ROLE_USER)
async def dispatch():
"""Send the task and return only once it settles.
This coroutine never touches cancellation -- it's a plain dispatch
that would, left alone, run for 100+ steps. Wrapped in
asyncio.ensure_future() below so it keeps running in the
background while the rest of this notebook goes looking for it.
"""
async with open_client() as client:
request = SendMessageRequest(message=message)
last = None
async for chunk in client.send_message(request):
last = chunk
return last
send_future = asyncio.ensure_future(dispatch())
Step 2 — Find the In-Flight Task¶
The task is running somewhere on the server right now, but this
notebook doesn't have its id yet. LLMAgentA2AExecutor publishes only
a single terminal status, so there's no incremental chunk to read one
off the way a streaming executor would allow. Instead, a fresh client
polls list_tasks(), which the SDK backs with the same task store
execute() persists to immediately on submission, well before the
agent's first LLM call returns, so the task shows up as
TASK_STATE_SUBMITTED almost immediately, then TASK_STATE_WORKING
once the agent starts.
list_tasks() returns every in-flight task on the server, not just
this notebook's own; if the server was already running from a
previous cell or a prior run, a stale or concurrent task could match
just as well. Each listed task's history carries the original
dispatch message back, so matching on message.message_id (captured
in Step 1) picks out this notebook's own task specifically, not merely
the first pending one found.
task_id = None
try:
async with open_client() as observer:
for _ in range(50):
await asyncio.sleep(0.2)
listing = await observer.list_tasks(
ListTasksRequest(history_length=1),
)
for task in listing.tasks:
state = TaskState.Name(task.status.state)
is_pending = state in (
"TASK_STATE_SUBMITTED",
"TASK_STATE_WORKING",
)
is_ours = any(
m.message_id == message.message_id for m in task.history
)
if is_pending and is_ours:
task_id = task.id
if task_id:
break
if task_id is None:
send_future.cancel()
msg = "Task never appeared in list_tasks()."
raise RuntimeError(msg)
except Exception:
stop_server()
raise
print(f"Found in-flight task {task_id}")
Found in-flight task 0a9af83b-868c-4ef5-967c-8584fc9e11b2
Step 3 — Cancel It¶
Now that the task id is known, a second fresh client (not the one
still waiting on dispatch()) asks the server to cancel it.
try:
async with open_client() as observer:
cancel_result = await observer.cancel_task(
CancelTaskRequest(id=task_id),
)
except Exception:
stop_server()
raise
print(
"cancel_task() response state:",
TaskState.Name(cancel_result.status.state),
)
cancel_task() response state: TASK_STATE_CANCELED
Step 4 — Confirm It Actually Stopped¶
Cancelling settles the original dispatch() call too, once the
server-side work actually tears down. Its own final chunk should agree
with what cancel_task() already reported.
try:
final_chunk = await send_future
except Exception:
stop_server()
raise
final_state = TaskState.Name(final_chunk.task.status.state)
print("send_message()'s own final chunk state:", final_state)
send_message()'s own final chunk state: TASK_STATE_CANCELED
hailstone_server.log is the more convincing evidence: the log
goes quiet right after the first "Processing Step" line, with no
further tool call ever logged. The agent's own LLM call was still in
flight when cancel() tore down the background task underneath it.
with open("hailstone_server.log") as f:
print(f.read())
✓ Using Ollama Cloud (kimi-k2.7-code:cloud) INFO: Started server process [2556248] INFO: Waiting for application startup. INFO: Application startup complete. INFO: Uvicorn running on http://0.0.0.0:9300 (Press CTRL+C to quit) INFO: 127.0.0.1:58726 - "GET /.well-known/agent-card.json HTTP/1.1" 200 OK INFO: 127.0.0.1:58738 - "GET /.well-known/agent-card.json HTTP/1.1" 200 OK INFO (llm_agents_fs.LLMAgent) : 🚀 Starting task: Compute the full hailstone sequence starting at 27, until it reaches 1. INFO (llm_agents_fs.TaskHandler) : ⚙️ Processing Step: Compute the full hailstone sequence starting at 27, until it reaches 1. INFO: 127.0.0.1:58764 - "POST / HTTP/1.1" 200 OK INFO: 127.0.0.1:58776 - "POST / HTTP/1.1" 200 OK INFO: 127.0.0.1:58748 - "POST / HTTP/1.1" 200 OK
Cleanup¶
stop_server()