Explore models
Try any live model in the browser, then copy the call your AI agent makes through the Summer MCP server.
Video 16
Gemini Omni Flash
Fast video with synchronized audio from text, a still image, references, or a motion clip.
Veo 3.1 Fast
The faster, cheaper Veo tier for drafting shots with native audio.
Veo 3.1
Google's premium Veo tier for coherent shots with native audio.
ByteDance
Seedance 2.0 Fast
Fast multimodal video with audio and clips up to 15 seconds.
ByteDance
Seedance 2.5
Cinematic long takes with sound, from text, a still image or a set of references.
ByteDance
Seedance 2.0
The higher-quality Seedance tier for audio-backed clips up to 15 seconds.
Kuaishou
Kling 3.0
Fluid motion, native audio, and motion transfer from a reference clip onto your character.
MiniMax
MiniMax H3 Max
Ranked first for image to video, with sharp motion up to 15 seconds.
Alibaba
Wan 3.0
Alibaba's Wan 3.0, smooth motion and native audio from a prompt, a frame or references.
Alibaba
HappyHorse 1.1
Native audio with multilingual lip sync, written straight into the prompt.
Veo 3.1 Lite
The lowest-cost Veo tier, 720p clips with native audio.
xAI
Grok Imagine Video 1.5
xAI's video model with sound, from a prompt, a still or references.
Vidu
Vidu Q3 Pro
Vidu's Q3 Pro model, clips with audio up to 16 seconds.
PixVerse
PixVerse V6
Fast, low-cost clips with sound.
Lightricks
LTX-2.5 Pro
Lightricks' open-source audio and video model, sound and picture in one pass.
Luma
Luma Ray 3.2
Luma's cinematic video model, rich camera work and seamless loops.
Image 10
xAI
Grok Imagine 2
Concept images from xAI's Grok Imagine 2.0 at standard 1K output.
Nano Banana 2 Lite
Fast 1K image generation and reference editing for clean game-asset references.
ByteDance
Seedream 5.0 Pro
Strong prompt understanding, multilingual text, and structured layouts.
Nano Banana Lite
Fast 1K generation and reference editing for asset iteration.
xAI
Grok Imagine
Fast, sharp concept art and textures at a low price per image.
OpenAI
GPT Image 2.5
OpenAI's top-ranked image model, with real light, exact text and faithful edits.
OpenAI
GPT Image 2
Precise text, UI, and structured image generation from OpenAI.
Meta
Muse Image
Fast, near-frontier images and precise edits at the lowest price.
Gemini 2.5 Flash Image
Fast general-purpose image generation and editing from Google.
Black Forest Labs
FLUX.2
The model behind Summer Pixel: crisp, consistent pixel art sprites.
3D 10
Tencent
Hunyuan 3D 3.1 Pro
High-quality geometry and textures from Tencent's Hunyuan 3D v3.1.
Tencent
Hunyuan 3D 3.1 Rapid
Fast, lower-cost 3D drafts for shape checks.
Hyper3D
Rodin 2.5
Production-focused geometry, PBR materials, and clean topology from up to five references.
Meshy
Meshy 6
Meshy 3D with the rigging and animation path for characters.
Tripo
Tripo H3.1
Detailed geometry, PBR, quad output, and multi-view generation.
Microsoft
Trellis 2
Image to 3D mesh through the Summer MCP, from any engine.
Meshy
Meshy 7.1
Meshy's newest model, with quad or triangle meshes and one-step rigging for characters.
Tripo
Tripo P2
The first model to build native quad meshes, with texture tiers up to extreme.
Hitem3D
Hi3D V3.0
The finest geometry detail, built at 2048 cubed voxel resolution.
Hyper3D
Rodin 2.5 Fast
Hyper3D's fast tier: light meshes up to 20K faces for quick drafts.
Audio 14
ElevenLabs
Eleven v3
ElevenLabs' most emotionally rich speech model, with multi-speaker dialogue.
ElevenLabs
Eleven Multilingual v2
Stable, lifelike speech for longer narration in many languages.
ElevenLabs
Eleven Flash v2.5
Low-latency speech for fast iteration.
ElevenLabs
ElevenLabs Text to Sound v2
Sound effects, loops, and ambience from a prompt.
ElevenLabs
Eleven Music v2
Music with stronger vocals, arrangement, and long-form structure.
ElevenLabs
Eleven Music v1
Studio-grade music from a prompt or a composition plan.
Gemini 3.8 Flash TTS
Google's newest speech model, directed in plain words.
Gemini 3.8 Flash Lite TTS
The lighter Gemini 3.8 speech model, for more lines for less.
MiniMax
MiniMax Speech 2.8 HD
Studio-grade speech with emotion, pitch and speed controls.
xAI
xAI TTS
Grok's expressive voices, with laugh, sigh and pause tags.
Alibaba
Qwen Audio 3.0 TTS
Alibaba's speech model with 45 character voices.
Inworld
Inworld TTS-1.5 Max
Game-ready character voices from Inworld, at the lowest price here.
ElevenLabs
ElevenLabs Scribe v2
Transcripts with word timings and speaker labels.
OpenAI
Whisper Large v3
OpenAI's Whisper in 99 languages, or translated to English.