Gemini 3.5 Flash-Lite

Fast high-throughput Gemini model for low-latency agentic and document-heavy workloads

Available via OpenAI-compatible API

Gemini 3.5 Flash-Lite Overview

Gemini 3.5 Flash-Lite is a lightweight multimodal model in Google's Gemini 3.5 family, built for high-throughput and latency-sensitive production workloads. It is well suited to agentic search, document processing, coding subtasks, routing, summarization, and other flows where fast response times and scale matter as much as model quality. It combines the Lite tier's efficiency profile with stronger long-context, coding, and agent-style performance than earlier Flash-Lite generations.

Token based
Input tokens / 1M$0.30
Output tokens / 1M$2.50
Cached input / 1M$0.03

Commercial use