Skip to main content
Agnes 3.0 Flash Max is a paid Agnes AI text model available on the international site. Use agnes-3.0-flash-max as the model name. Regular input, cached input, and output tokens are billed separately.

Overview

Use the following endpoints to integrate Agnes 3.0 Flash Max into chat, coding, and tool-calling applications.

Chat Completions API

Endpoint and headers

Request fields

Agnes 3.0 Flash Max supports the following request fields.

Image URL input

Pass public image URLs together with text in messages[].content.

Basic request

Response format

Chat Completions responses use an OpenAI-compatible structure:

Tool-calling request

Read generated text from choices[].message.content. For a tool call, inspect choices[].message.tool_calls, execute the requested function in your application, append its result to messages, and submit the next Chat Completions request.

Responses API

The Responses API accepts text or structured messages through input.
Read generated text from a message item in output where output[].type is message and output[].content[].type is output_text. The response object may also contain usage, error, and incomplete_details.

Messages API

Agnes 3.0 Flash Max supports the Anthropic-compatible Messages API.
Read generated text from content[] items where content[].type is text.

Thinking mode

Enable Thinking mode when your integration needs more deliberate task decomposition or reasoning.

Best practices

State the task objective, repository or runtime context, constraints, expected output, and tool permissions. Return tool results to the conversation before asking the model for the next action.
Write narrow tool descriptions and JSON schemas. Validate tool arguments in your application before executing side-effecting actions.
Break complex work into verifiable stages, and preserve the objective, constraints, and key tool results across each execution round.
Model availability, rate limits, and billing are determined by your Agnes AI account and API key.

Pricing and billing

Prices below are in USD for the international site. M means one million tokens. This is a paid model; the Agnes 3.0 Flash free offer does not apply. Cached-input pricing applies only to input tokens confirmed as cache hits by the service. Other input tokens use the regular input price; the same input token is not billed twice. Availability and rate limits depend on account entitlements.