> ## Documentation Index
> Fetch the complete documentation index at: https://wiki.agnes-ai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Agnes 3.0 Flash Max

> Agnes 3.0 Flash Max 付费文本模型的 API 接入指南、请求与响应示例及美元定价。

<Info>
  Agnes 3.0 Flash Max 是 Agnes AI 的付费文本模型，在国际站提供服务。模型名称为 `agnes-3.0-flash-max`，按普通输入、缓存输入和输出 Token 分别计费。
</Info>

## 概述

通过以下接口接入 Agnes 3.0 Flash Max，构建对话、编码与工具调用应用。

| 项目 | 内容 |
| - | - |
| 模型名称 | `agnes-3.0-flash-max` |
| Base URL | `https://apihub.agnes-ai.com/v1` |
| 接口 | `POST /v1/chat/completions`、`POST /v1/responses`、`POST /v1/messages` |
| 计费方式 | 付费，按 Token 用量计费 |
| 服务区域 | 国际站 |

## Chat Completions API

### Endpoint 与请求头

```text theme={null}
POST https://apihub.agnes-ai.com/v1/chat/completions
```

```bash theme={null}
-H "Authorization: Bearer YOUR_API_KEY"
-H "Content-Type: application/json"
```

### 请求字段

Agnes 3.0 Flash Max 支持以下请求字段。

| 字段 | 类型 | 必填 | 说明 |
| - | - | - | - |
| `model` | string | 是 | 使用 `agnes-3.0-flash-max`。 |
| `messages` | array | 是 | 对话消息数组，包含 `system`、`user` 和 `assistant` 角色。 |
| `messages[].content` | string / array | 是 | 纯文本，或包含 `text`、`image_url` 的内容块数组。 |
| `temperature` | number | 否 | 控制输出随机性。 |
| `top_p` | number | 否 | 控制核采样。 |
| `max_tokens` | integer | 否 | 最大输出 Token 数。 |
| `stream` | boolean | 否 | 设为 `true` 时返回流式响应。 |
| `tools` | array | 否 | 用于函数调用工作流的工具定义。 |
| `tool_choice` | string / object | 否 | 控制模型是否使用工具以及如何使用。 |
| `chat_template_kwargs` | object | 否 | Thinking 等兼容扩展能力的扩展字段。 |

## 图像 URL 输入

在 `messages[].content` 中，将公开可访问的图像 URL 与文本一起传入。

```json theme={null}
{
  "role": "user",
  "content": [
    {
      "type": "text",
      "text": "请描述这张图片中的关键信息。"
    },
    {
      "type": "image_url",
      "image_url": {
        "url": "https://example.com/image.jpg"
      }
    }
  ]
}
```

### 基础请求

```bash theme={null}
curl https://apihub.agnes-ai.com/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "agnes-3.0-flash-max",
    "messages": [
      {
        "role": "user",
        "content": "请说明智能体应如何选择并调用工具。"
      }
    ],
    "max_tokens": 1024
  }'
```

### 响应格式

Chat Completions 响应使用 OpenAI 兼容结构：

```json theme={null}
{
  "id": "chatcmpl_xxx",
  "object": "chat.completion",
  "model": "agnes-3.0-flash-max",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "智能体应根据任务和工具声明的能力选择工具。"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 24,
    "completion_tokens": 18,
    "total_tokens": 42
  }
}
```

### 工具调用请求

```json theme={null}
{
  "model": "agnes-3.0-flash-max",
  "messages": [
    {
      "role": "user",
      "content": "上海现在的天气怎么样？"
    }
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "获取指定城市当前天气。",
        "parameters": {
          "type": "object",
          "properties": {
            "city": { "type": "string" }
          },
          "required": ["city"]
        }
      }
    }
  ]
}
```

从 `choices[].message.content` 读取生成文本。模型返回工具调用时，读取 `choices[].message.tool_calls`，在你的应用中执行相应函数，将结果追加到 `messages` 后发起下一次 Chat Completions 请求。

## Responses API

Responses API 使用 `input` 传入文本或结构化消息。

```text theme={null}
POST https://apihub.agnes-ai.com/v1/responses
```

```bash theme={null}
curl https://apihub.agnes-ai.com/v1/responses \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "agnes-3.0-flash-max",
    "input": "总结这个任务，并列出完成它需要的工具。",
    "max_output_tokens": 1024
  }'
```

| 字段 | 类型 | 必填 | 说明 |
| - | - | - | - |
| `model` | string | 是 | 使用 `agnes-3.0-flash-max`。 |
| `input` | string / array | 是 | 纯文本提示词或结构化输入消息。 |
| `max_output_tokens` | integer | 否 | 最大输出预算。 |

从 `output` 中类型为 `message` 的项读取生成内容，其中 `output[].content[].type` 为 `output_text`。响应对象还可能包含 `usage`、`error` 和 `incomplete_details`。

## Messages API

Agnes 3.0 Flash Max 支持 Anthropic 兼容 Messages API。

```text theme={null}
POST https://apihub.agnes-ai.com/v1/messages
```

```bash theme={null}
curl https://apihub.agnes-ai.com/v1/messages \
  -H "x-api-key: YOUR_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "agnes-3.0-flash-max",
    "max_tokens": 1024,
    "messages": [
      {
        "role": "user",
        "content": "请起草一份简洁的产品公告，并列出后续行动。"
      }
    ]
  }'
```

| 字段 | 类型 | 必填 | 说明 |
| - | - | - | - |
| `model` | string | 是 | 使用 `agnes-3.0-flash-max`。 |
| `max_tokens` | integer | 是 | 最大输出 Token 预算。 |
| `messages` | array | 是 | 含 `user` 与 `assistant` 角色的消息数组。 |
| `messages[].content` | string / array | 是 | 纯文本或 Anthropic 兼容的内容块。 |
| `system` | string / array | 否 | 系统指令。 |
| `temperature` | number | 否 | 控制输出随机性。 |
| `stream` | boolean | 否 | 设为 `true` 时返回流式响应。 |

从 `content[]` 中 `content[].type` 为 `text` 的内容块读取生成文本。

## Thinking 模式

需要更充分的任务拆解或推理时，可启用 Thinking 模式。

<Tabs>
  <Tab title="OpenAI 兼容格式">
    ```json theme={null}
    {
      "model": "agnes-3.0-flash-max",
      "messages": [
        {
          "role": "user",
          "content": "请规划此仓库任务的实现步骤。"
        }
      ],
      "chat_template_kwargs": {
        "enable_thinking": true
      }
    }
    ```
  </Tab>

  <Tab title="Anthropic 兼容格式">
    ```json theme={null}
    {
      "model": "agnes-3.0-flash-max",
      "max_tokens": 2048,
      "messages": [
        {
          "role": "user",
          "content": "请规划此仓库任务的实现步骤。"
        }
      ],
      "thinking": {
        "type": "enabled",
        "budget_tokens": 2048
      }
    }
    ```
  </Tab>
</Tabs>

## 最佳实践

<AccordionGroup>
  <Accordion title="Agnes Code 与智能体任务">
    明确任务目标、仓库或运行环境上下文、限制条件、预期输出和工具权限。执行完工具后，应将工具结果返回对话，再请求模型给出下一步动作。
  </Accordion>

  <Accordion title="工具调用">
    工具描述与 JSON Schema 应保持精确、聚焦。在应用侧执行会产生副作用的操作前，应先校验工具参数。
  </Accordion>

  <Accordion title="长任务执行">
    将复杂任务拆分为可验证的阶段，并在每轮执行中保留目标、限制条件和关键工具结果。
  </Accordion>
</AccordionGroup>

<Note>
  模型可用性、速率限制和计费以你的 Agnes AI 账户与 API Key 权限为准。
</Note>

## 价格与计费

以下为国际站美元价格，`M` 表示一百万 Token。本模型为付费模型，不适用 Agnes 3.0 Flash 的免费优惠。

| 计费项 | 单价 |
| - | -: |
| 普通输入（Prefill input） | `$0.08 / M tokens` |
| 缓存输入（Cached input） | `$0.008 / M tokens` |
| 输出（Output） | `$0.40 / M tokens` |

缓存价格仅适用于服务确认命中缓存的输入 Token；其余输入按普通输入单价计费，同一输入 Token 不重复计费。模型可用性与速率限制以账户权限为准。


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.