> ## Documentation Index
> Fetch the complete documentation index at: https://wiki.agnes-ai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Agnes 3.0 Flash Max

> API integration, request and response examples, and USD pricing for the paid Agnes 3.0 Flash Max text model.

<Info>
  Agnes 3.0 Flash Max is a paid Agnes AI text model available on the international site. Use `agnes-3.0-flash-max` as the model name. Regular input, cached input, and output tokens are billed separately.
</Info>

## Overview

Use the following endpoints to integrate Agnes 3.0 Flash Max into chat, coding, and tool-calling applications.

| Item | Value |
| - | - |
| Model | `agnes-3.0-flash-max` |
| Base URL | `https://apihub.agnes-ai.com/v1` |
| Endpoints | `POST /v1/chat/completions`, `POST /v1/responses`, `POST /v1/messages` |
| Billing | Paid, based on token usage |
| Availability | International site |

## Chat Completions API

### Endpoint and headers

```text theme={null}
POST https://apihub.agnes-ai.com/v1/chat/completions
```

```bash theme={null}
-H "Authorization: Bearer YOUR_API_KEY"
-H "Content-Type: application/json"
```

### Request fields

Agnes 3.0 Flash Max supports the following request fields.

| Field | Type | Required | Description |
| - | - | - | - |
| `model` | string | Yes | Use `agnes-3.0-flash-max`. |
| `messages` | array | Yes | Conversation messages with `system`, `user`, and `assistant` roles. |
| `messages[].content` | string / array | Yes | Plain text or content blocks containing `text` and `image_url`. |
| `temperature` | number | No | Controls output randomness. |
| `top_p` | number | No | Controls nucleus sampling. |
| `max_tokens` | integer | No | Maximum number of output tokens. |
| `stream` | boolean | No | Returns a streamed response when `true`. |
| `tools` | array | No | Tool definitions for function-calling workflows. |
| `tool_choice` | string / object | No | Controls whether and how the model calls tools. |
| `chat_template_kwargs` | object | No | Extension field for Thinking and other compatible features. |

## Image URL input

Pass public image URLs together with text in `messages[].content`.

```json theme={null}
{
  "role": "user",
  "content": [
    {
      "type": "text",
      "text": "Describe the key information in this image."
    },
    {
      "type": "image_url",
      "image_url": {
        "url": "https://example.com/image.jpg"
      }
    }
  ]
}
```

### Basic request

```bash theme={null}
curl https://apihub.agnes-ai.com/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "agnes-3.0-flash-max",
    "messages": [
      {
        "role": "user",
        "content": "Explain how an agent should choose and call a tool."
      }
    ],
    "max_tokens": 1024
  }'
```

### Response format

Chat Completions responses use an OpenAI-compatible structure:

```json theme={null}
{
  "id": "chatcmpl_xxx",
  "object": "chat.completion",
  "model": "agnes-3.0-flash-max",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "An agent should select a tool based on the task and the tool's declared capability."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 24,
    "completion_tokens": 18,
    "total_tokens": 42
  }
}
```

### Tool-calling request

```json theme={null}
{
  "model": "agnes-3.0-flash-max",
  "messages": [
    {
      "role": "user",
      "content": "What is the weather in Shanghai?"
    }
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get the current weather for a city.",
        "parameters": {
          "type": "object",
          "properties": {
            "city": { "type": "string" }
          },
          "required": ["city"]
        }
      }
    }
  ]
}
```

Read generated text from `choices[].message.content`. For a tool call, inspect `choices[].message.tool_calls`, execute the requested function in your application, append its result to `messages`, and submit the next Chat Completions request.

## Responses API

The Responses API accepts text or structured messages through `input`.

```text theme={null}
POST https://apihub.agnes-ai.com/v1/responses
```

```bash theme={null}
curl https://apihub.agnes-ai.com/v1/responses \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "agnes-3.0-flash-max",
    "input": "Summarize this task and list the tools you would need.",
    "max_output_tokens": 1024
  }'
```

| Field | Type | Required | Description |
| - | - | - | - |
| `model` | string | Yes | Use `agnes-3.0-flash-max`. |
| `input` | string / array | Yes | Plain-text prompt or structured input messages. |
| `max_output_tokens` | integer | No | Maximum output budget. |

Read generated text from a message item in `output` where `output[].type` is `message` and `output[].content[].type` is `output_text`. The response object may also contain `usage`, `error`, and `incomplete_details`.

## Messages API

Agnes 3.0 Flash Max supports the Anthropic-compatible Messages API.

```text theme={null}
POST https://apihub.agnes-ai.com/v1/messages
```

```bash theme={null}
curl https://apihub.agnes-ai.com/v1/messages \
  -H "x-api-key: YOUR_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "agnes-3.0-flash-max",
    "max_tokens": 1024,
    "messages": [
      {
        "role": "user",
        "content": "Draft a concise product announcement and list the next actions."
      }
    ]
  }'
```

| Field | Type | Required | Description |
| - | - | - | - |
| `model` | string | Yes | Use `agnes-3.0-flash-max`. |
| `max_tokens` | integer | Yes | Maximum output-token budget. |
| `messages` | array | Yes | Messages with `user` and `assistant` roles. |
| `messages[].content` | string / array | Yes | Plain text or Anthropic-compatible content blocks. |
| `system` | string / array | No | System instruction. |
| `temperature` | number | No | Controls output randomness. |
| `stream` | boolean | No | Returns a streamed response when `true`. |

Read generated text from `content[]` items where `content[].type` is `text`.

## Thinking mode

Enable Thinking mode when your integration needs more deliberate task decomposition or reasoning.

<Tabs>
  <Tab title="OpenAI-compatible format">
    ```json theme={null}
    {
      "model": "agnes-3.0-flash-max",
      "messages": [
        {
          "role": "user",
          "content": "Plan the implementation steps for this repository task."
        }
      ],
      "chat_template_kwargs": {
        "enable_thinking": true
      }
    }
    ```
  </Tab>

  <Tab title="Anthropic-compatible format">
    ```json theme={null}
    {
      "model": "agnes-3.0-flash-max",
      "max_tokens": 2048,
      "messages": [
        {
          "role": "user",
          "content": "Plan the implementation steps for this repository task."
        }
      ],
      "thinking": {
        "type": "enabled",
        "budget_tokens": 2048
      }
    }
    ```
  </Tab>
</Tabs>

## Best practices

<AccordionGroup>
  <Accordion title="Agnes Code and agent tasks">
    State the task objective, repository or runtime context, constraints, expected output, and tool permissions. Return tool results to the conversation before asking the model for the next action.
  </Accordion>

  <Accordion title="Tool calling">
    Write narrow tool descriptions and JSON schemas. Validate tool arguments in your application before executing side-effecting actions.
  </Accordion>

  <Accordion title="Long-running tasks">
    Break complex work into verifiable stages, and preserve the objective, constraints, and key tool results across each execution round.
  </Accordion>
</AccordionGroup>

<Note>
  Model availability, rate limits, and billing are determined by your Agnes AI account and API key.
</Note>

## Pricing and billing

Prices below are in USD for the international site. `M` means one million tokens. This is a paid model; the Agnes 3.0 Flash free offer does not apply.

| Billing item | Price |
| - | -: |
| Prefill input | `$0.08 / M tokens` |
| Cached input | `$0.008 / M tokens` |
| Output | `$0.40 / M tokens` |

Cached-input pricing applies only to input tokens confirmed as cache hits by the service. Other input tokens use the regular input price; the same input token is not billed twice. Availability and rate limits depend on account entitlements.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.