Muse Spark 1.1 Pricing, Rate Limits, and API Costs (2026)
In this article
Muse Spark 1.1 is Meta's latest multimodal model, priced at $1.25 per million input tokens, $4.50 per million output tokens, and $0.15 per million cached input tokens. It sits in the mid-range API tier, making it a practical choice for production workloads that need capable reasoning without flagship-model costs.
Muse Spark 1.1 pricing breakdown
According to the official Meta AI developer documentation, Spark 1.1 uses a standard input/output token pricing model with a separate cached input tier:
| Token type | Price per 1M tokens |
|---|---|
| Input tokens | $1.25 |
| Output tokens | $4.50 |
| Cached input tokens | $0.15 |
Cached input is charged at a heavily discounted rate ($0.15 vs $1.25), so any application with repeated context (system prompts, document retrieval, tool definitions) should use prompt caching to cut costs significantly. A workflow that passes a 10,000-token system prompt on every request saves roughly 88% on those input tokens by caching.
What is Muse Spark 1.1?
Muse Spark 1.1 is the latest generation of Meta's Spark model family, released as part of the Muse platform. As covered in the Meta developer blog, Spark 1.1 is designed for developers building creative and multimodal applications, with improvements in instruction following, longer context handling, and output quality over the original Spark release.
Discussion on Hacker News highlights that Spark 1.1 is generating interest among developers evaluating it against OpenAI and Anthropic alternatives, particularly for tasks where cached context is expensive at scale. The consensus is that the $0.15 cached input rate is one of the more competitive offerings in its class.
Muse Spark 1.1 rate limits
Rate limits for Spark 1.1 are defined per API tier in the Meta dev docs. Limits apply across requests per minute (RPM), tokens per minute (TPM), and tokens per day (TPD). Higher tiers unlock higher throughput after spend thresholds are met.
- Requests per minute (RPM): Varies by tier; free tier starts conservatively to prevent abuse
- Tokens per minute (TPM): Scales with tier; production accounts get substantially higher limits
- Tokens per day (TPD): Daily hard cap applied across all models in the Muse platform
If you hit a rate limit, the API returns a 429 Too Many Requests response. Meta recommends exponential backoff with jitter as the standard retry strategy. Unlike some providers, Spark 1.1 does not have per-message or per-session caps; limits are purely token and request based.
How rate limits compare to similar models
| Model | Input (per 1M) | Output (per 1M) | Cached input |
|---|---|---|---|
| Muse Spark 1.1 | $1.25 | $4.50 | $0.15 |
| Claude Haiku 3.5 | $0.80 | $4.00 | $0.08 |
| GPT-4o mini | $0.15 | $0.60 | $0.075 |
| OpenAI Codex | Varies | Varies | N/A |
Spark 1.1 sits above GPT-4o mini but below Claude Sonnet and GPT-4o on price. Whether the output quality justifies the cost depends on your specific use case.
Free tier and getting started
Meta provides a free tier for Muse Spark 1.1 access through the developer portal. Free tier accounts have reduced rate limits and a monthly token cap. This is sufficient for prototyping but not for production traffic.
To move off the free tier, you add a payment method to your Meta AI developer account. Paid tier limits activate automatically once billing is configured. There is no manual upgrade step or support ticket required.
Cost estimation for common workloads
Using the $1.25 input / $4.50 output / $0.15 cached input pricing:
- 1,000 API calls with 500 input + 500 output tokens each: ~$2.88 total ($0.625 input + $2.25 output)
- Same calls with 400 tokens cached, 100 uncached input: ~$2.37 (saving ~18% vs no caching)
- High-volume: 10M input tokens / 5M output tokens per month: $12.50 input + $22.50 output = $35/month
For agentic applications with long system prompts, caching can reduce costs by 50% or more. The math favors heavy prompt caching whenever your system prompt exceeds ~2,000 tokens.
Key takeaways
- Spark 1.1 is priced at $1.25/M input, $4.50/M output, $0.15/M cached input
- Use prompt caching for any repeated context - the 88% discount on cached tokens is significant
- Rate limits are tiered; free tier is for prototyping only
- No per-message caps - all limits are token and request based
- 429 errors mean you've hit a rate limit; use exponential backoff
If you're also working with Claude Code and want to monitor your own AI tool usage in real time, Usagebar tracks your Claude Code 5-hour window, weekly cap, and context usage from the macOS menu bar.
Related reading
- OpenAI Codex pricing and rate limits
- Kiro pricing and free tier
- Claude Cowork pricing and limits
- Claude Code vs Cursor pricing comparison
- Codex usage limit and reset time
Sources
Never Get Locked Out Mid-Task Again
Never hit your usage limits unexpectedly. Usagebar lives in your menu bar and shows your 5-hour and weekly limits at a glance.
Get Usagebar$9 — one-time, lifetime updates