ModelMeta
820 model profiles·28 providers·Refresh cadence hourly
Back to models
Cohere

Rerank v4.0 Fast

Active

Fast reranking model optimized for latency

Context window

33K

32,768 tokens

Max output

Not published

No output limit in registry

Input price

$0.20

Per 1M input tokens

Output price

Not published

No official price captured yet

Modalities

Text

Input: Text; output: Text

API surface

Not mapped

Endpoint support not captured

Overview

Where this model fits best

Use this section to quickly decide whether the model belongs in chat, coding, reasoning, embedding, rerank, vision, audio, or agent workflows.

Use cases

What this model should be considered for

Selection signal
rerank

Best fit

Use this model when you need a well-documented, structured option inside the registry and want a single place to inspect pricing, capabilities, and operational limits.

Pricing

Pricing and billing signals

Known prices are shown per 1M tokens. Missing official prices are marked as not published instead of being treated as free.

Standard API

Input

$0.20 / 1M

Output

Not published

Default request pricing when using the primary endpoint.

Controls

Runtime knobs worth knowing

Supported request parameters help developers understand sampling, output limits, reasoning controls, and structured output behavior.

Temperature

1

0 to 2

Top P

1

0 to 1

Presence Penalty

0

-2 to 2

Frequency Penalty

0

-2 to 2

Specifications

Technical reference

Canonical identifiers, family, modalities, token limits, training information, and update metadata.

Model ID

rerank-v4.0-fast

Use this exact identifier in API calls and SDK configuration.

Provider

Cohere

Access type

Closed

Input modalities

text

Output modalities

text