xAI Grok chat, image, video and voice models
Explore/xAI Models
xAI

xAI API models

xAI Grok chat, image, video and voice models

xAI's Grok ecosystem spans chat, image, video and voice. PoYo.ai lists Grok Imagine for visual generation and xAI TTS 1 for expressive multilingual text-to-speech.

xAI Models API - pricing and performance

Run xAI models through PoYo.ai with transparent API pricing, model pages, and a consistent async generation workflow.

Transparent pricing

Each model page lists its current PoYo.ai API price and billing unit.

Model comparison

Compare task types, categories, and related providers before integration.

Async API workflow

Submit generation jobs and retrieve results through PoYo.ai model APIs.

ModelCategoryTask typesPrice
Grok 4.6

grok-4.6

ChatChat$1.60Per 1M input
Grok 4.7

grok-4.7

ChatChat$1.60Per 1M input
Grok Imagine Image 2.0

grok-imagine-image-2.0

ImageText to Image, Image to Image$0.04Per generation
xAI TTS 1

xai-tts-1

MusicText to Speech$0.012Per generation
Grok Imagine

grok-imagine-image

VideoText to Image, Image to Video, Uncensored$0.03Per generation
Grok Imagine Video 1.5

grok-imagine-video-1.5

VideoImage to Video$0.072Per second
Grok Imagine Image Quality

grok-imagine-image-quality

ImageText to Image, Image to Image$0.04Per generation

Frequently asked questions

What xAI models are available on PoYo.ai?+

PoYo.ai lists Grok Imagine entries for image and video generation plus xAI TTS 1 for text-to-speech. xAI's broader official API lineup also includes Grok chat models.

Does xAI support chat, image, video and voice?+

Yes. Official xAI documentation covers Grok chat models, Grok Imagine image and video models, and voice capabilities including text-to-speech. PoYo.ai exposes visual generation and xAI TTS 1.

Can I compare xAI with other providers?+

Yes. Compare xAI with OpenAI and Google for multimodal coverage, with Seedream and Black Forest Labs for image generation, with Kling, Seedance, Wan, Runway and Google Veo for video generation, and with ElevenLabs or Google TTS for voice.