Rerank v4.0 Fast
Fast reranking model optimized for latency
Context window
32,768 tokens
Max output
No output limit in registry
Input price
Per 1M input tokens
Output price
No official price captured yet
Modalities
Input: Text; output: Text
API surface
Endpoint support not captured
Overview
Where this model fits best
Use this section to quickly decide whether the model belongs in chat, coding, reasoning, embedding, rerank, vision, audio, or agent workflows.
Use cases
What this model should be considered for
Best fit
Use this model when you need a well-documented, structured option inside the registry and want a single place to inspect pricing, capabilities, and operational limits.
Pricing
Pricing and billing signals
Known prices are shown per 1M tokens. Missing official prices are marked as not published instead of being treated as free.
Standard API
Input
Output
Default request pricing when using the primary endpoint.
Controls
Runtime knobs worth knowing
Supported request parameters help developers understand sampling, output limits, reasoning controls, and structured output behavior.
Temperature
0 to 2
Top P
0 to 1
Presence Penalty
-2 to 2
Frequency Penalty
-2 to 2
Specifications
Technical reference
Canonical identifiers, family, modalities, token limits, training information, and update metadata.
Model ID
rerank-v4.0-fastUse this exact identifier in API calls and SDK configuration.
Provider
Access type
Input modalities
Output modalities