Vision Small
Fast, low-cost multimodal model for understanding text, images, audio, video, and PDFs, with tool calling and a 1M-token context window.
ReasoningTool CallingAttachmentsStructured Output
Context Window
1M
Max Output
66K
Temperature
Yes
Open Weights
No
Knowledge Cutoff
N/A
Released
2024-05-15
Last Updated
2026-09
Modalities
Input:TextImageAudioVideoPDF
→Output:Text
Available Providers (1)
| Provider | Input /1M | Output /1M | Cache Read /1M | Cache Write /1M | Reasoning | Status |
|---|---|---|---|---|---|---|
| Rp 18.775 | Rp 56.501 | — | — | ✓ | — |