> For the complete documentation index, see [llms.txt](https://docs.delphi.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.delphi.ai/api-immortal-only/experience-agent.md).

# Run a Delphi in your harness

Use the OpenAI Responses API with your own tools, streaming, and Pydantic AI.

Connect your harness to a Delphi through the OpenAI Responses API. Delphi supplies the owner's identity and freshly retrieved knowledge. Your harness supplies the task, executes function tools, and keeps the conversation history.

## Connect

| Setting   | Value                                   |
| --------- | --------------------------------------- |
| Base URL  | `https://api.delphi.ai/v4/agents`       |
| Model     | `delphi`                                |
| API key   | A Delphi key with `generate:text` scope |
| Transport | OpenAI Responses API                    |

Create a key in [API Keys](https://www.delphi.ai/~/actions/keys). The key selects its owner's Delphi automatically. No slug is needed. Calls use PUBLIC knowledge and anonymous visitor context; caller-supplied identity or access-tier headers cannot unlock private knowledge or visitor memory.

The OpenAI client's `Authorization: Bearer` header works directly. You can also use `x-api-key`; if you send both, they must contain the same key. Keep your key on your server and set it as the `DELPHI_API_KEY` environment variable.

Model discovery is available at `GET /v4/agents/models`. Choose the Responses transport in your harness; Chat Completions and Anthropic Messages are not supported on this base URL.

## Make a request

These examples were verified with OpenAI `2.40.0` and Pydantic AI `1.105.0`:

```bash
python -m pip install "openai==2.40.0" "pydantic-ai-slim[openai]==1.105.0"
```

Save each Python example in its own file and run it with `DELPHI_API_KEY` set.

```python
import os

from openai import OpenAI

with OpenAI(
    api_key=os.environ["DELPHI_API_KEY"],
    base_url="https://api.delphi.ai/v4/agents",
    timeout=180,
    max_retries=0,
) as client:
    response = client.responses.create(
        model="delphi",
        input="How would you decide which project to pursue next?",
        store=False,
    )
    if response.status != "completed":
        raise RuntimeError(f"Response ended with {response.status}")
    print(response.output_text)
```

## Use tools and continue the conversation

This Pydantic AI example runs a local inventory function, requests a structured recommendation, then asks a follow-up using the complete history. Replace the inventory fixture with your own service call. The fixture's prices come from your tool; the recommendation also draws on the Delphi's knowledge.

```python
import asyncio
import os
from dataclasses import dataclass, field
from decimal import Decimal
from typing import Annotated

from openai import AsyncOpenAI
from pydantic import BaseModel, Field, WithJsonSchema
from pydantic_ai import Agent, NativeOutput, RunContext
from pydantic_ai.models.openai import OpenAIResponsesModel, OpenAIResponsesModelSettings
from pydantic_ai.providers.openai import OpenAIProvider
from pydantic_ai.usage import UsageLimits


class Inventory(BaseModel):
    sku: str
    quantity: int
    unit_price: Decimal


class Recommendation(BaseModel):
    owner_principle: str = Field(description="A principle found in the Delphi's knowledge")
    source: str = Field(description="Its supporting source, or say if unavailable")
    quantity: int
    unit_price: Annotated[Decimal, WithJsonSchema({"type": "number"})]
    total_cost: Annotated[Decimal, WithJsonSchema({"type": "number"})]
    recommendation: str


@dataclass
class HarnessTools:
    executions: list[Inventory] = field(default_factory=list)


async def lookup_inventory(ctx: RunContext[HarnessTools], sku: str) -> Inventory:
    """Look up current stock and price for a workshop item."""
    inventory = Inventory(sku=sku, quantity=7, unit_price=Decimal("13.25"))
    ctx.deps.executions.append(inventory)
    return inventory


async def main() -> None:
    async with AsyncOpenAI(
        api_key=os.environ["DELPHI_API_KEY"],
        base_url="https://api.delphi.ai/v4/agents",
        timeout=180,
        max_retries=0,
    ) as client:
        agent = Agent(
            OpenAIResponsesModel("delphi", provider=OpenAIProvider(openai_client=client)),
            deps_type=HarnessTools,
            tools=[lookup_inventory],
            output_type=NativeOutput(Recommendation),
            model_settings=OpenAIResponsesModelSettings(openai_store=False),
            instructions=(
                "For the initial request, look up SKU REPAIR-KIT with the inventory tool. "
                "Combine its quantity and price with a relevant principle from your knowledge "
                "to recommend whether to buy all available units. Calculate the total cost. "
                "Identify the principle and its source honestly; never invent a principle "
                "or attribute inventory data to the owner. Do not look up inventory again "
                "during a follow-up unless explicitly asked to refresh it."
            ),
        )
        tools = HarnessTools()
        result = await agent.run(
            "Should I buy repair kits for my workshop?",
            deps=tools,
            usage_limits=UsageLimits(request_limit=4),
        )
        print(result.output.model_dump_json(indent=2))

        followup = await agent.run(
            "Using the same inventory result, explain why that principle applies. "
            "Do not refresh inventory.",
            deps=tools,
            message_history=result.all_messages(),
            usage_limits=UsageLimits(request_limit=2),
        )
        print(followup.output.model_dump_json(indent=2))
        print(f"Inventory lookups: {len(tools.executions)}")


asyncio.run(main())
```

`WithJsonSchema` keeps the price fields numeric in the submitted JSON schema while Python validates them as `Decimal`. The inventory fixture supplies seven units at 13.25 each, for a total of 92.75. The follow-up reuses that result.

If you manage the tool loop yourself:

1. Send function definitions in `tools` and your conversation in `input`.
2. Append the returned `response.output` items to your history. Execute each `function_call` in your harness, then append a `function_call_output` with its original `call_id` and the tool's result in `output`.
3. Send the full history again with the same tools and instructions. Include results for every call in a parallel batch before adding another message.

Preserve function-call items and argument strings; replacing them with prose loses the tool contract. Delphi rebuilds its grounding on every request, including tool continuations.

## Stream text

Set `stream=True` to receive Responses events. Text chunks are provisional until the stream reaches `response.completed`.

```python
import os

from openai import OpenAI

with OpenAI(
    api_key=os.environ["DELPHI_API_KEY"],
    base_url="https://api.delphi.ai/v4/agents",
    timeout=180,
    max_retries=0,
) as client:
    completed = False
    with client.responses.create(
        model="delphi",
        input="Explain one principle you use when making difficult decisions.",
        store=False,
        stream=True,
    ) as events:
        for event in events:
            if event.type == "response.output_text.delta":
                print(event.delta, end="", flush=True)
            elif event.type == "response.completed":
                completed = True
            elif event.type in ("response.incomplete", "response.failed"):
                raise RuntimeError(f"Response ended with {event.response.status}")
    if not completed:
        raise RuntimeError("The stream ended before completion")
    print()
```

Streaming also carries function-call events. With Pydantic AI, use `agent.run_stream(...)`; call `await result.get_output()` before using a structured result, and retain `result.all_messages()` for the next turn.

## Supported options and limits

* **History:** text messages, function calls, and function results. Use `input`, not `messages`. Set `store:false` and supply complete history on every call. No Delphi threads or visitor memories are created.
* **Tools:** caller-defined functions, parallel calls, and `tool_choice` values `auto`, `none`, `required`, or a named function. Names must be unique; `build_context` is reserved. JSON schema references must be local fragments.
* **Output:** text, JSON objects, or JSON schemas through `text.format`. Incomplete responses can contain unfinished JSON; check status before parsing. Truncated function calls must not be executed.
* **Generation:** `max_output_tokens`, `temperature`, and `top_p` are supported. The output-token budget includes the internal context call and any structured output repair. Disconnecting cancels generation.
* **Unsupported:** `previous_response_id`, saved conversations, hosted tools, images, audio, and background execution. Unsupported options return `400`.

Your key's request limits and the owner's daily generation budget apply. Each Responses request, including a tool continuation or retry, counts as a generation attempt. The generation budget is shared across the owner's keys. Model discovery does not consume a generation attempt.

## Handle failures

HTTP errors use `{"error":{"message":"...","type":"...","code":"..."}}`.

| Status | What to check                                                                      |
| ------ | ---------------------------------------------------------------------------------- |
| `400`  | Unsupported options, malformed tool history, or conflicting auth headers.          |
| `401`  | Missing or invalid Delphi API key.                                                 |
| `403`  | API access, `generate:text` scope, or a required tool blocked by a safety finding. |
| `404`  | The key owner's Delphi is missing or inactive.                                     |
| `429`  | Request limits or daily generation budget; respect `Retry-After` when present.     |
| `5xx`  | A service could not complete the request.                                          |

After streaming starts, failures may arrive as `response.failed`; token limits or content filtering can produce `response.incomplete`. Check the terminal event as well as the HTTP status.

The examples disable automatic HTTP retries. Retrying starts a new generation; responses are not stored or replayed by idempotency key. Your harness controls whether to repeat a tool action that may already have completed.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the following URL with the `ask` and `goal` query parameters:

```
GET https://docs.delphi.ai/api-immortal-only/experience-agent.md?ask=<question>&goal=<user_goal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is what the user is ultimately trying to achieve, the reason they need the answer. Sharing it helps GitBook give you a better, more relevant answer. A goal is most helpful when it describes the outcome the user wants rather than restating the question. For example, with `ask=how do I create an API token`, a goal like `build a script that syncs our docs to a CMS` lets GitBook tailor the answer to that use case.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
