Vision Medium
Balanced multimodal model pairing a 1M-token context window with deeper reasoning for analysis, content creation, and tool use across text, image, audio, video, and PDF inputs.
ReasoningTool CallingAttachmentsStructured Output
Context Window
1M
Max Output
66K
Temperature
Yes
Open Weights
No
Knowledge Cutoff
N/A
Released
2024-05-15
Last Updated
2026-09
Modalities
Input:TextImageAudioVideoPDF
→Output:Text
Available Providers (1)
| Provider | Input /1M | Output /1M | Cache Read /1M | Cache Write /1M | Reasoning | Status |
|---|---|---|---|---|---|---|
| Rp 75.275 | Rp 225.825 | — | — | ✓ | — |