
AI APIs for developers
AI Chat API
Build writing, reasoning and coding features with the PoYo AI Chat API. Compare LLM models, conversation protocols, supported inputs, response controls and token pricing.
Matching models
27 models

$2.28 1M input
kimi-k3
Moonshot AI's 2.8T-parameter flagship multimodal model with a 1M-token context window, built for long-horizon coding, deep reasoning, knowledge work, and tool-using agents.

$0.85 1M input
claude-sonnet-5
Claude Sonnet 5 API access for agentic coding, long-running agents, browser and computer use, professional workflows, 1M context, and up to 128K output.

$0.112 1M input
deepseek-v4-flash
DeepSeek V4 Flash is a fast, economical DeepSeek V4 API model with 1M-context chat, 284B total parameters, and 13B active parameters.

$0.348 1M input
deepseek-v4-pro
DeepSeek V4 Pro is a stronger reasoning, coding, and agent workflow model with 1M-context chat, 1.6T total parameters, and 49B active parameters.

$8.00 1M input
gpt-6-astra
GPT-6 Astra API for advanced reasoning, coding, and agent workflows. Use Chat Completions or Responses with pay-as-you-go pricing on PoYo.

$1.60 1M input
gpt-6-sol
GPT-6 Sol API on PoYo. Use Chat Completions or Responses at $1.6 input and $8 output per million tokens, 20% below official standard rates.

$8.00 1M input
claude-fable-5-1
Anthropic's most capable generally available model for long-horizon coding, knowledge work, research, vision, 1M context, and up to 128K output.

$0.06 1M input
gemini-3.7-flash
Google's GA Flash model for agentic coding, multimodal reasoning, web development, long-context knowledge work, and tool-using applications.

$0.06 1M input
gemini-3.8-flash
Google's best reasoning and coding Flash model for long-horizon software engineering, autonomous agents, and enterprise workflows.

$0.056 1M input
gpt-5-6-luna
GPT-5.6 API access for Sol, Terra, and Luna through the Responses API, from high-throughput automation to frontier coding and agent workflows.

$0.08 1M input
gpt-6-luna
GPT-6 Luna API on PoYo. Use Chat Completions or Responses at $0.08 input and $0.4 output per million tokens, 20% below official standard rates.

$1.60 1M input
grok-4.6
xAI's frontier model for long-running agents, coding, engineering, knowledge work, and ambitious interactive or visual projects.

$1.60 1M input
grok-4.7
Grok 4.7 API on PoYo. Use Chat Completions at $1.6 input and $4.8 output per million tokens, 20% below official standard rates.

$2.00 1M input
claude-opus-5
Anthropic's flagship agentic model for advanced coding, long-horizon execution, computer use, deep research, and professional knowledge work.

$3.20 1M input
claude-opus-5-5
Claude Opus 5.5 API on PoYo. Use Chat Completions or Claude Messages at $3.2 input and $16 output per million tokens, 20% below official standard rates.

$1.60 1M input
claude-sonnet-5-5
Claude Sonnet 5.5 API on PoYo. Use Chat Completions or Claude Messages at $1.6 input and $8 output per million tokens, 20% below official standard rates.

$0.30 1M input
deepseek-flash
DeepSeek V4.1 Flash is a 552B MoE chat model with native vision, 1M context, 384K max output, and 8B/16B active parameters for cheaper agent workloads.

$1.60 1M input
gpt-6.1-sol
GPT-6.1 Sol API on PoYo. Use Chat Completions or Responses at $1.6 input and $8 output per million tokens, 20% below official standard rates.

$3.00 1M input
gpt-5.5
GPT-5.5 API access for chat, coding, reasoning, and production agent workflows.

$1.05 1M input
gpt-5.4
GPT-5.4 API access for chat, coding, reasoning, and production agent workflows.

$0.90 1M input
gemini-3.5-flash
Gemini 3.5 Flash API access for chat, coding, reasoning, and production agent workflows.

$0.438 1M input
gpt-5.2
GPT-5.2 API access for chat, coding, reasoning, and production agent workflows.

$4.00 1M input
claude-opus-4-8
Claude Opus 4.8 supports long-context chat, agentic coding, professional reasoning, high-output workflows, 1M context, and up to 128K output.

$4.00 1M input
claude-opus-4-7
Anthropic's most capable generally available model for frontier coding, deep reasoning, and high-resolution vision.

$1.44 1M input
claude-sonnet-4-6
Claude Sonnet 4.6 and Claude Opus 4.6 API access for coding, agents, computer use, and long-context reasoning.

$0.80 1M input
claude-opus-4-5-20251101
Anthropic's Claude 4.5 series model — Opus, Sonnet & Haiku

$0.40 1M input
gemini-3-flash-preview
Google's Gemini 3 series model — Flash Preview & Pro Preview
Chat Model APIs - Pricing and Model Fit
Compare providers, supported inputs, output formats and prices before choosing a model.
Exact task match
Models appear here only when their catalog metadata includes Chat.
Provider comparison
Review model families across providers without leaving the task directory.
API-ready paths
Each model card links to PoYo pricing, examples, playground controls and API details.
| Model | Provider | Task types | Price |
|---|---|---|---|
| Kimi K3 kimi-k3 | Moonshot AI | Chat | $2.281M input |
| Claude Sonnet 5 claude-sonnet-5 | Anthropic | Chat | $0.851M input |
| DeepSeek V4 Flash deepseek-v4-flash | DeepSeek | Chat | $0.1121M input |
| DeepSeek V4 Pro deepseek-v4-pro | DeepSeek | Chat | $0.3481M input |
| GPT-6 Astra gpt-6-astra | OpenAI | Chat | $8.001M input |
| GPT-6 Sol gpt-6-sol | OpenAI | Chat | $1.601M input |
| Claude Fable 5.1 claude-fable-5-1 | Anthropic | Chat | $8.001M input |
| Gemini 3.7 Flash gemini-3.7-flash | Chat | $0.061M input | |
| Gemini 3.8 Flash gemini-3.8-flash | Chat | $0.061M input | |
| GPT-5.6 gpt-5-6-luna | OpenAI | Chat | $0.0561M input |
| GPT-6 Luna gpt-6-luna | OpenAI | Chat | $0.081M input |
| Grok 4.6 grok-4.6 | xAI | Chat | $1.601M input |
| Grok 4.7 grok-4.7 | xAI | Chat | $1.601M input |
| Claude Opus 5 claude-opus-5 | Anthropic | Chat | $2.001M input |
| Claude Opus 5.5 claude-opus-5-5 | Anthropic | Chat | $3.201M input |
| Claude Sonnet 5.5 claude-sonnet-5-5 | Anthropic | Chat | $1.601M input |
| DeepSeek V4.1 Flash deepseek-flash | DeepSeek | Chat | $0.301M input |
| GPT-6.1 Sol gpt-6.1-sol | OpenAI | Chat | $1.601M input |
| GPT-5.5 gpt-5.5 | OpenAI | Chat | $3.001M input |
| GPT-5.4 gpt-5.4 | OpenAI | Chat | $1.051M input |
| Gemini 3.5 Flash gemini-3.5-flash | Chat | $0.901M input | |
| GPT-5.2 gpt-5.2 | OpenAI | Chat | $0.4381M input |
| Claude Opus 4.8 claude-opus-4-8 | Anthropic | Chat | $4.001M input |
| Claude Opus 4.7 claude-opus-4-7 | Anthropic | Chat | $4.001M input |
| Claude 4.6 API claude-sonnet-4-6 | Anthropic | Chat | $1.441M input |
| Claude 4.5 Series claude-opus-4-5-20251101 | Anthropic | Chat | $0.801M input |
| Gemini 3 Series gemini-3-flash-preview | Chat | $0.401M input |
Frequently asked questions
What is an AI chat API?
+
An AI chat API accepts conversation messages and returns a language-model response. It enables writing, reasoning, coding and assistant features in your own application. Available input types and controls depend on the selected model and API protocol.
Which LLM API model should I choose?
+
Evaluate models with the prompts, conversation lengths and outputs your product uses. Compare answer quality, instruction following, speed and token cost. Include difficult and unsuccessful examples in the evaluation rather than selecting from one impressive response.
Does every PoYo chat model use the same request format?
+
No. Use the model documentation to choose OpenAI-compatible Chat Completions, Claude Messages or Gemini native format. Message structures, authentication details and supported controls can differ. Adapt the request to the documented protocol instead of changing only the model name.
How do I preserve multi-turn conversation context?
+
Store the relevant messages in your application and send the history required by the selected protocol. Keep within the model context limit and remove or summarize material that no longer helps the task. A new API request should not be assumed to remember earlier requests automatically.
Does the AI chat API support streaming responses?
+
Check streaming support and the request options for the chosen model and endpoint. If available, handle partial events, completion and errors in your client. Keep a non-streaming path when your workflow needs a complete response before validation or downstream processing.
Can I send images or other media to a chat model?
+
Only use media types and message structures documented for that model. Some LLM endpoints are text-only and others support multimodal inputs. Check size, format and content limitations before implementing uploads, and test how well the model handles your actual media.
Can I request JSON or another structured response?
+
Some models expose structured-output controls; others can only be guided by prompts. Confirm the supported mechanism, supply a clear schema where available and validate the returned data in your application. Do not treat a request for JSON as proof that every response will satisfy your schema.
Can a chat model browse the web or call tools automatically?
+
Only when the endpoint supports the capability and your application provides or enables the required tools. A model can describe a tool call without executing it. Check the documented tool protocol and keep execution, permissions and result handling explicit in your integration.
How can I reduce inaccurate AI chat answers?
+
Provide relevant source material, make the task precise and ask the model to distinguish missing information from known facts. Validate important outputs against trusted data and test recurring failure cases. Model choice and prompting can improve quality, but generated answers still need appropriate review.
How is AI chat API pricing calculated?
+
Chat pricing commonly separates input and output tokens, and some models have additional tier distinctions. Read the selected model pricing and usage response. Include conversation history in your estimate, since repeatedly sending long context can increase the cost of each turn.
What should I do about context errors, rate limits or timeouts?
+
Check the returned error and the selected model limits. Reduce oversized context, control concurrent requests and use bounded retries where appropriate. For a partial streaming response, record what was received and decide how your application should recover instead of silently duplicating a conversation turn.
Where can I find AI chat API integration examples?
+
Open the model page and its API documentation for the correct endpoint, authentication, message body and examples. The model-specific llms.txt resource also identifies the protocol. Keep keys on the server and validate outputs before using them in automated business actions.