Skip to main content
Agnes 2.5 Flash is a generally available language model upgraded from Agnes 2.0 Flash. It keeps the same OpenAI-compatible Chat Completions integration path while improving the experience for coding, agent workflows, tool calling, multi-turn conversations, reasoning, and image understanding.

Model

agnes-2.5-flash

API Endpoints

Chat Completions: POST /v1/chat/completions
Responses: POST /v1/responses
Messages: POST /v1/messages

Release Status

Generally available to users with Agnes API access.

Upgrade Path

API-compatible upgrade from the deprecated agnes-2.0-flash model.

Overview

Agnes 2.5 Flash is designed as the next-step upgrade for developers already using Agnes 2.0 Flash. In most integrations, you only need to replace the model value with agnes-2.5-flash; the Base URL, endpoint, headers, message format, streaming format, tool-calling format, and image URL input format remain the same. The model focuses on a smoother developer experience, stronger instruction following, more stable multi-turn output, and improved code-specific capability for generation, debugging, refactoring, explanation, and agentic coding workflows.
Use agnes-2.5-flash directly as the model name. agnes-2.0-flash is deprecated, so existing integrations should migrate as soon as possible.

Core Capabilities

Chat Completions

Generate high-quality responses for conversations, applications, and business systems.

Multi-turn Conversations

Maintain context consistency across continuous interactions.

Image URL Input

Accept visual content through publicly accessible image URLs.

Image Understanding

Analyze screenshots, describe images, answer visual questions, and extract visual information.

Tool Calling

Support function calling and external tool orchestration.

Agent Workflows

Improved planning, execution, context tracking, and multi-step task completion.

Code-specialized Tasks

Optimized for code generation, debugging, refactoring, explanation, and patch-style development workflows.

Streaming

Return responses in real time for a better interactive experience.

Use Cases

AI Assistants

General Q&A, productivity assistants, personal assistants, and in-app copilots.

Autonomous Agents

Multi-step task execution, planning, tool use, and workflow scheduling.

Coding Assistants

Code generation, bug fixing, refactoring suggestions, code review, test generation, and code explanation.

Customer Support

FAQ automation, support chatbots, and service workflow automation.

Search and Q&A

Retrieval-based answers, summarization, and information extraction.

Image Understanding

Screenshot analysis, image description, visual Q&A, and structured extraction.

Upgrade from Agnes 2.0 Flash

If you already call agnes-2.0-flash, the 2.5 Flash migration is intentionally small.
For existing integrations, migration can usually be handled as a model-name change. Do not continue using the deprecated agnes-2.0-flash model as a compatibility fallback.

API Reference

Endpoint

Headers

Request Parameters

Image URL Input

Agnes 2.5 Flash supports passing text and image URLs in the same messages request.

Request Examples

Response Format

Response Fields

Responses API

In addition to Chat Completions, this model supports the OpenAI Responses API. Use input instead of messages.

Responses endpoint

Responses request parameters

Responses output format

The current response does not include a top-level output_text convenience field. Extract generated text from message items where output[].type is message and output[].content[].type is output_text.
Reasoning items are optional and can use either content[].reasoning_text or summary[].summary_text. Token usage field names can also vary by model: support both input_tokens / output_tokens and prompt_tokens / completion_tokens.
If status is incomplete, inspect incomplete_details and retry with a larger max_output_tokens value. Reasoning models can consume part of the output budget before producing assistant text.

Messages API

This model also supports the Anthropic-compatible Messages API. Send conversation input in messages and authenticate with x-api-key.

Messages endpoint

Messages headers

Messages request parameters

Messages request example

Messages response format

Read generated text from content blocks where content[].type is text. If stop_reason is max_tokens, retry with a larger max_tokens value.

Thinking Mode

For coding, debugging, reasoning, and agent workflows, you can enable Thinking mode to improve task decomposition and problem-solving quality.
For regular coding tasks, start with budget_tokens: 2048. For complex debugging, refactoring, or multi-step agent workflows, increase the budget as needed.

Best Practices

Limits and Pricing

Agnes 2.5 Flash is generally available. Availability, rate limits, and billing behavior follow the entitlement shown for your Agnes AI account and API key.

Integration Checklist

Use agnes-2.5-flash as the model name.
Basic chat completion requests must include model and messages.
Image inputs must use publicly accessible image_url values.
Set stream to true when you need streaming responses.