Client
Configure the Python client, handle errors, and generate chat, text, and structured output.
BlazingAgents is the entry point for sync code and AsyncBlazingAgents for async code. You configure the key, timeouts, and headers once, then call resources and generation methods from it. Creating a client makes no network request, but it raises ValueError when it finds no API key.
Construct a client
Signatures:
BlazingAgents(
*,
api_key: str | None = None,
base_url: str = "https://api.blazingagents.com",
timeout: Timeout = 60.0,
default_headers: Mapping[str, str] | None = None,
http_client: httpx.Client | None = None,
on_response: Callable[[ResponseObservation], None] | None = None,
)
AsyncBlazingAgents(
*,
api_key: str | None = None,
base_url: str = "https://api.blazingagents.com",
timeout: Timeout = 60.0,
default_headers: Mapping[str, str] | None = None,
http_client: httpx.AsyncClient | None = None,
on_response: Callable[[ResponseObservation], None] | None = None,
)An explicit api_key wins over BLAZING_AGENTS_API_KEY. default_headers go on every request, and the extra_headers and timeout arguments on a single call override them for that call.
The SDK owns two headers. It always sets Authorization from your tenant API key, overwriting any value you pass. It strips X-Request-Id from default_headers and extra_headers, and rejects an injected HTTPX client whose default headers contain it. To correlate requests with your own IDs, use client_request_id instead.
The SDK closes the HTTPX client it creates when you close the SDK client. An http_client you pass in stays yours; the SDK never closes it.
Ordinary requests time out after 60 seconds by default, and None disables the timeout. Streaming requests keep the connect, write, and pool timeouts but have no read deadline. The SDK never retries on its own.
agent()
Signature: agent(agent_id: str) -> AgentClient
Returns resources scoped to one agent, without a network request. Use client.agent(agent_id).skills for every operation on that agent's skills. On AsyncBlazingAgents, agent() returns an AsyncAgentClient with the same skills member; await its request methods and use async for with its iterator. See Skills.
with_options()
Signature: with_options(*, client_request_id: str) -> BlazingAgents
Returns a copy of the client that sends X-Client-Request-Id on every resource and generation call. The original client is unchanged. On AsyncBlazingAgents the method is also synchronous and returns an AsyncBlazingAgents.
correlated = client.with_options(client_request_id="checkout-attempt-42")
agent = correlated.agents.get("ag_0123456789abcdef")Generation methods also accept client_request_id per call, and every method accepts extra_headers and timeout.
close()
Signature: close() -> None
Closes the connections a sync client owns. Prefer with BlazingAgents(...) as client, which calls close() on exit. An injected HTTPX client stays open.
aclose()
Signature: await aclose() -> None
Closes the connections an AsyncBlazingAgents owns. Prefer async with AsyncBlazingAgents(...) as client, which awaits aclose() on exit. The async client never runs sync calls through the event loop.
Response observation and request IDs
Pass on_response to see every response the client receives, including API errors, malformed bodies, and stream handshakes. The callback gets one ResponseObservation per response:
| Field | Meaning |
|---|---|
method | HTTP method |
path | Path without query parameters |
status | HTTP status |
duration_ms | Elapsed request time in milliseconds |
request_id | Server request ID from X-Request-Id, when present |
client_request_id | Your correlation ID, when present |
The SDK ignores exceptions your callback raises. A connection failure with no response does not call it.
Response models expose the server request ID as _request_id, which is not serialized. Generated text and stream objects expose it as request_id. Keep it when you contact support. Each retry you make is a new attempt with a new request ID.
Errors
Every SDK exception derives from BlazingAgentsError.
| Exception | Raised when |
|---|---|
APIStatusError | The server returns an error status. Carries status_code, headers, the server code, details, param, request_id, a safe response_body, and retry_after |
APIConnectionError | The network fails before a complete response |
APITimeoutError | An HTTP timeout fires; also an APIConnectionError |
StreamError | Reading, ownership, headers, decoding, or finalization fails after stream headers arrive. Carries status_code, headers, request_id, and retry_after |
ObjectTruncationError | A complete response contains incomplete JSON |
ObjectJSONDecodeError | A complete response contains invalid JSON |
ObjectValidationError | Decoded JSON does not match the requested output type |
from blazing_agents import APIConnectionError, APIStatusError, APITimeoutError
try:
agent = client.agents.get("ag_0123456789abcdef")
except APIStatusError as error:
if error.code == "not_found":
print(error.request_id, error.retry_after)
except APITimeoutError:
...
except APIConnectionError:
...str(error) for an APIStatusError starts with the code in brackets, such as [model_validation_unavailable] Provider model discovery is unavailable, so logs show the code without extra work. error.code holds the bare code, such as model_validation_unavailable.
Branch on error.code, not the message. Cancellation, KeyboardInterrupt, and SystemExit pass through unwrapped. Decide whether to retry from the operation, the status, and retry_after.
Logging and telemetry
The SDK is silent by default. It sends no telemetry and starts no background tasks. If you enable the blazing_agents logger at debug level, each record contains only the method, the path without query, the status, elapsed time, and the request ID. It never logs credentials, headers, query values, bodies, schemas, file data, or stream content.
Generation methods
The client has five generation methods: chat() for conversations that Blazing Agents stores as sessions, plus buffered and streaming forms of stateless text and structured output. Every call runs one metered turn. All arguments are keyword-only.
Give each call exactly one input: a literal message or prompt, or a saved prompt through prompt_id. variables works only with prompt_id. Every generation method also accepts version to pin an agent version, user_id and metadata to attribute the turn to an end user, client_request_id, extra_headers, and timeout.
AsyncBlazingAgents has the same five method names; you await them.
Methods
chat()
Signature: chat(*, agent_id, message=..., prompt_id=..., variables=..., trigger=..., message_id=..., session_id=..., version=..., user_id=..., metadata=..., client_request_id=None, extra_headers=None, timeout=...) -> ChatStream
Sends a message in a session and returns a ChatStream of the server's AI SDK SSE bytes, exactly as sent. The SDK does not decode or re-encode UIMessageChunk values.
with client.chat(
agent_id="ag_0123456789abcdef",
message={
"id": "message-1",
"role": "user",
"parts": [{"type": "text", "text": "Hello"}],
},
) as stream:
session_id = stream.session_id
for chunk in stream:
print(chunk.decode(), end="")Omit session_id to start a new session. Its ss_... ID is available as stream.session_id before you read the body. Pass it on a later call to continue the conversation. You can pin version only when you start a session, not when you continue one. trigger="regenerate-message" works only in an existing session and can target a message_id.
With the async client, call stream = await client.chat(...), then use async with stream and async for chunk in stream.
completion()
Signature: completion(*, agent_id, prompt=..., prompt_id=..., variables=..., version=..., user_id=..., metadata=..., client_request_id=None, extra_headers=None, timeout=...) -> Completion
Returns the complete text answer. Completion is a str subclass that also carries request_id.
result = client.completion(
agent_id="ag_0123456789abcdef",
prompt="Summarize this request.",
)
print(str(result), result.request_id)Use await client.completion(...) with AsyncBlazingAgents.
completion_stream()
Signature: completion_stream(*, agent_id, prompt=..., prompt_id=..., variables=..., version=..., user_id=..., metadata=..., client_request_id=None, extra_headers=None, timeout=...) -> CompletionStream
Streams the text answer as decoded text deltas. get_final_text() reads anything you have not consumed yet and returns the full Completion. Reading to the end closes the stream; call close() to stop early.
with client.completion_stream(
agent_id="ag_0123456789abcdef",
prompt="Write a release note.",
) as stream:
for delta in stream:
print(delta, end="")
final = stream.get_final_text()With the async client, call stream = await client.completion_stream(...), then use async with, async for, await stream.get_final_text(), and await stream.aclose() to stop early.
object()
Signature: object(*, agent_id, output_type=..., json_schema=..., prompt=..., prompt_id=..., variables=..., version=..., user_id=..., metadata=..., client_request_id=None, extra_headers=None, timeout=...) -> T | JsonValue
Returns structured output. Pass exactly one of output_type or json_schema. With a Pydantic-compatible output_type, the SDK derives the JSON Schema, decodes the complete response, and validates it with Pydantic's TypeAdapter, so you get an instance of output_type. With a raw json_schema, you get the decoded JSON value without validation.
from pydantic import BaseModel
class Summary(BaseModel):
title: str
risks: list[str]
summary = client.object(
agent_id="ag_0123456789abcdef",
prompt="Summarize the release.",
output_type=Summary,
)Use await client.object(...) with the async client.
object_stream()
Signature: object_stream(*, agent_id, output_type=..., json_schema=..., prompt=..., prompt_id=..., variables=..., version=..., user_id=..., metadata=..., client_request_id=None, extra_headers=None, timeout=...) -> ObjectStream[T] | ObjectStream[JsonValue]
Streams structured output as raw JSON text deltas. It never yields partial Pydantic models. get_final_object() reads any remaining data and validates only after the stream finishes. Invalid, truncated, or mismatched output raises the matching Object...Error.
with client.object_stream(
agent_id="ag_0123456789abcdef",
prompt="Summarize the release.",
output_type=Summary,
) as stream:
for json_delta in stream:
print(json_delta, end="")
summary = stream.get_final_object()With the async client, call stream = await client.object_stream(...), then use async with, async for, await stream.get_final_object(), and await stream.aclose() to stop early.
Stream ownership and failures
Each chat, completion, and object stream has exactly one reader. Iterating a second time or after close raises StreamError, as do a failed read, a malformed session location, or an incomplete finish. Object streams raise the more specific object errors. Reading to the end closes the stream. To stop early, use a context manager; closing an active stream closes its HTTP connection.
Status, connection, and timeout failures before the stream starts raise the normal client exceptions. Stream objects expose the status, headers, and request_id before you read them.