Mistral Medium 3.5
- API model IDs
- mistral-medium-2604
- Context window
- 256,000 tokens
- Max output
- 256,000 tokens
- Released
- 2026-04-28
- Lifecycle
- active
- License
- open-weights
Lifecycle verified 2026-09-08 (0 days ago). Source
Data handling
- Trains on API data by default
- Yes
- Retention
- Not published
- Zero data retention
- on-request
- Residency
- eu
- HIPAA BAA
- Not available
- SOC 2 / FedRAMP
- SOC 2 yes, FedRAMP Not published
Mistral's help center states that using API calls to improve its services is on by default, with customers retaining the right to opt out at any time from the Admin panel's Privacy menu; the Vibe and API opt-out toggles are separate. Zero data retention is available on paid plans for supported stateless API calls but is not a self-service toggle: it requires a written request with a stated reason, which Mistral reviews and may approve or deny, after which it appears in Admin privacy settings. Data is hosted in the EU by default, with an explicit opt-in to a US API endpoint available. No specific retention period in days for API prompts and responses is stated on a page this audit could confirm, so it is recorded as unknown rather than the widely repeated but unverifiable 30-day figure. Mistral states it complies with SOC 2 Type II and ISO 27001/27701; a HIPAA Business Associate Agreement and any FedRAMP authorization are not confirmed on a page this audit could read, so they are recorded conservatively rather than assumed.
Verified 2026-09-08 (0 days ago). Source 1, Source 2, Source 3, Source 4
Pricing snapshot
$1.5 input, $7.5 output, per million tokens.
This is a snapshot verified on 2026-09-08, 0 days ago, not a live price. Confirm on the provider’s pricing page before you budget against it.
Where you can call it
- direct · mistral-medium-2604 First-party Mistral La Plateforme API.
- self-host Released as open weights under a Modified MIT license.
Mistral's API reference (docs.mistral.ai/api) does not publish a max output token figure separate from context window: it states only that prompt tokens plus max_tokens must not exceed the model's context length, so the output cap recorded here equals the published context window.