specs at a glance

Leaderboard rank#9 of 30
SWE-bench Verified92.2%
SWE-bench Pro—
Terminal-Bench82.0% (TB2.1, vendor)
Input price / 1M—
Output price / 1M—
Context window—
Open weightsNo
AccessAPI (Fireworks Serverless research preview)
MakerFireworks AI (built on Moonshot Kimi K3)

how good is Ember-1 at coding?

Ember-1 sits at #9 of 30 ranked models, posting 92.2% on SWE-bench Verified — 4.8 points behind #1 Claude Opus 5. Terminal-Bench (agentic terminal work): 82.0% (TB2.1, vendor).

Score provenance: Vendor-reported (Fireworks, Sept 23 2026, Fireworks harness): SWE-bench Verified 92.2% using 15.5% fewer tokens than Kimi K3; Terminal Bench 2.1 82.0% (-51.9% tokens vs K3 Max); DeepSWE 1.1 75.2% (-23.7% tokens); live A/B tests 35% fewer tokens per task. No independent score: vals.ai archived its SWE-bench Verified board on 2026-09-05 and it does not list Ember-1 (checked 2026-09-28).

Its base, Kimi K3, measures 93.4% independently on vals.ai, so treat the 1.2-point gap as indicative only, since the two figures come from different harnesses. Pricing not stated in the announcement. See /p/fireworks-ember-1-kimi-k3-fewer-reasoning-tokens/.

where can you use it?

Available via API (Fireworks Serverless research preview). As a proprietary model, you're on the maker's infrastructure and release schedule.

Full storyFireworks Ember-1: Kimi K3 Quality on 40% Fewer Tokens

Primary source

Ranked on our AI Coding Leaderboard — scores confirmed against primary sources only, updated 2026-09-28.