Skip to main content
Agnes Video 2.5 is now available on the international site and uses an asynchronous video generation API. First call POST /v1/videos to create a task, then use the returned video_id with GET /agnesapi?video_id=<VIDEO_ID>&model_name=agnes-video-2.5 to retrieve progress and results. See “Billing” below for prices and billing rules.

Model ID

agnes-video-2.5

Create Task

POST /v1/videos

Retrieve Task

GET /agnesapi?video_id=<VIDEO_ID>&model_name=agnes-video-2.5

Pricing

720P: $0.025 / second; 960P: $0.040 / second; 2K: $0.055 / second.

Core Capabilities

Text-to-video

Generate videos with subject motion, scene dynamics, and camera movement from a text prompt.

First and Last Frame Control

Constrain the composition and transition with a first frame, a last frame, or both.

Multimodal References

Use images, audio, and videos as content, style, rhythm, or motion references.

Video-to-video Reference

Continue or reinterpret motion, visual appearance, and timing from a reference video.

Audio-visual Coordination

Use audio or a video soundtrack as a reference to improve rhythm and audio-visual consistency.

Multiple Aspect Ratios

Generate landscape, portrait, square, and ultrawide video outputs.

Quickstart

1. Get an API Key

Create an API key in the Agnes AI platform. Store and use the key only on your server. Never expose it in frontend code or a public repository.

2. Set the Base URL

International Base URL:
The examples below use environment variables:

3. Create a Video Task

Save video_id from the create response. id and task_id identify the asynchronous task, while video_id is used to retrieve progress and results.

4. Retrieve the Result

Use the query form with model_name for every mode. Poll every 1–2 seconds until status becomes completed or failed. When the task is complete, use metadata.url to play or download the video.

API Reference

Create a Video Task

Headers:

Common Request Parameters

Mode-specific Parameters

All media URLs must be publicly reachable by the Agnes AI service. Avoid URLs that require authentication, point to a private network, or expire before the task finishes.

Generation Mode Rules

keyframe attempts to preserve the input image as the actual first or last frame, making it suitable for start/end composition control. reference treats media as a content, style, motion, or rhythm reference and may recompose or retime the result.

Reference Video Objects

Each object in the videos array supports these fields: When require_audio is false, a reference video may omit audio. If it contains an audio track, the audio can also participate as a reference. When set to true, the source must contain an audio track or the request fails.

Request Examples

<Picture N>, <Audio N>, and <Video N> are numbered independently, starting from 1 in their respective arrays. For example, refer to the second item in images as <Picture 2>.

Create Response

Retrieve a Task

Completed response example:
Treat status and metadata.url as the source of truth. The URL is ready for delivery only when status is completed. In production, set a maximum polling duration and use backoff for network timeouts and 429 responses.

Python SDK Example

Pass mode, aspect_ratio, and media fields through extra_body; the SDK merges them into the top level of the request JSON.

Video Size and Aspect Ratio

size selects the output resolution tier and accepts "720P", "960P", or "2K". Use aspect_ratio to select the frame shape. WIDTHxHEIGHT and auto are not supported. The following table shows pixel examples for the 720P tier. The 960P and 2K tiers produce higher-resolution video in the selected aspect_ratio; use the API response as the source of truth for the actual width and height.

Parameter Restrictions

The following parameters and request patterns are not supported and return 400:
  • Passing reference videos through video_url, video_path, or video_reference; use videos[].url instead.
  • Passing media through input_reference or reference_url; use first_frame, last_frame, images, audios, or videos according to the selected mode.
  • Sending non-configurable fields such as width, height, fps, num_frames, quality, or num_inference_steps.
  • Passing pixel dimensions such as 1280x720 directly in size, or a value other than "720P", "960P", or "2K"; select the resolution tier with size and the frame shape with aspect_ratio.
  • Setting aspect_ratio to auto or a value outside the supported list.
  • Setting n to any value other than 1.
  • Using media fields that conflict with mode, or using reference without any reference media.

Error Handling

Failed task example:

Prompting Recommendations

For more consistent results, structure the prompt in this order:
  1. Subject and setting: Specify the people, objects, environment, and time.
  2. Action and change: Describe how the subject moves and how the scene evolves.
  3. Camera language: Specify push, pull, pan, tilt, tracking, fixed camera, or shot size.
  4. Visual style: Add lighting, color, material, realism, and atmosphere.
  5. Sound and rhythm: Describe ambient sound or action sounds, or reference an audio input.
  6. Consistency requirements: State which character, product, or composition details must remain unchanged.
In reference mode, explicitly name each media placeholder and its purpose, such as “Use <Picture 1> as the character reference and follow the rhythm of <Audio 1>.” This is more controllable than uploading media without explaining how it should be used.

Integration Checklist

  • Use the model ID agnes-video-2.5.
  • Use https://apihub.agnes-ai.com/v1 as the Base URL.
  • Save video_id from the create response; id and task_id identify the asynchronous task.
  • For every mode, use GET /agnesapi?video_id=<VIDEO_ID>&model_name=agnes-video-2.5 until the status is completed or failed. A video_id query without model_name is valid only for mode: "text".
  • Keep all media URLs publicly accessible until the task completes.
  • Pass seconds as a string from "4""12" and set n to 1.
  • Set size to "720P", "960P", or "2K" and use a supported aspect_ratio.
  • Never expose your API key in logs, client-side code, or public repositories.

Billing

Agnes Video 2.5 is billed by output resolution and output duration. The first 5 input images are free; each input image from the 6th onward costs $0.005.

Prices

Billing formula

The output resolution unit price is based on the requested size: 720P is $0.025 / second, 960P is $0.040 / second, and 2K is $0.055 / second.

Points billing

Points use the same metering structure as currency billing, but each points unit price differs from the corresponding currency amount:
Points use the same metering structure as currency billing. Refer to the Agnes AI platform for the points rate of each resolution tier and the excess-image points rate.

Billing example

Suppose you generate an 8-second 720P video with 7 input images:
The first 5 input images are free, so only 2 images incur an additional charge.