Google and Meta both shipped new coding-focused models this week
Google's Gemini 3.8 Flash and Meta's Muse Spark 1.3 launched a day apart, both aimed at coding and agentic tasks, and both priced as mid-tier workhorse models.
2 min read

By the numbers
- Gemini 3.8 Flash's per-million-token input/output price through 2026
- $0.75 / $3.75
- Muse Spark 1.3's per-million-token input/output price
- $1.25 / $4.25
- token context window both new models share
- 1M
Google shipped Gemini 3.8 Flash on September 2, and Meta followed with Muse Spark 1.3 within a day. Both target the same niche: mid-tier models priced for high-volume coding and agentic work, not just flagship benchmark scores.
Gemini 3.8 Flash: cheap, and built for long coding sessions
Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026, according to Google's own announcement. Google says it outperforms most larger frontier models on DeepSWE v1.1, a benchmark for solving long, complex software-engineering problems end to end.
On Terminal-Bench 2.1, a test of real command-line task completion, DataCamp reports Gemini 3.8 Flash scores 90.8%, up from 3.7 Flash's 81.6%. General reasoning, measured by Humanity's Last Exam, barely moved, 45.4% versus 45.7% for the prior version. Google is spending its improvement budget on coding and agentic skill, not raw knowledge.
Google also shipped a restricted Cyber variant through its Fairwind Program, limited to vetted defenders. Google reports it exceeds a 70% success rate on an internal vulnerability-discovery benchmark and scores 47.2% on CWE-Bench, a public benchmark for automatically patching known security flaws.
Muse Spark 1.3: a big jump, but not an open-weight model
Muse Spark 1.3 scores 61 on the Artificial Analysis Intelligence Index, per Trending Topics. That's close behind Claude Opus 5 and GPT-5.6 Sol, and ahead of the 59 the outlet credits to Gemini 3.8 Flash on the same scale. Meta's own reported gains are sharp: banking-task accuracy climbed from 35% to 47%, and terminal-operation accuracy from 80% to 85%.
Despite the framing some coverage has given it, Muse Spark 1.3 is not an open-weight release. LLM Stats describes it as "a proprietary multimodal reasoning model," available only through Meta's own API. Trending Topics notes Meta committed to releasing weights for the earlier Muse Spark 1.2, but hasn't decided whether 1.3's weights will follow. Pricing sits at $1.25 per million input tokens and $4.25 per million output tokens.
What this means for developers
If you're picking a model for an agent that writes and runs code, benchmark these two against each other on your own workload. Don't just compare them to the flagship models everyone already talks about; both are built and priced for exactly this job.
Don't assume Muse Spark 1.3 is something you can self-host. Unlike Meta's past Muse releases, this one is API-only for now, so any dependency on it is a dependency on Meta's infrastructure and pricing, not your own.
Watch the pattern here, not just the individual scores. Two major labs shipping mid-tier, coding-optimized models a day apart, at similar prices, is telling. It shows where the real competition has moved: not the top of the leaderboard, but the tier developers actually use for everyday, high-volume work.
Sources
Related articles

Thomson Reuters built its own frontier model for $40 million
Thomson Reuters says its in-house model Thomson matches frontier labs on legal work at a fraction of the cost. The benchmark footnotes deserve a closer read.

GPT-6 Astra scores 95% on one robot task, 10% on another
Robocurve ran GPT-6 Astra and Claude Fable 5.1 on real robot arms. Astra hit 19 of 20 on the easy task and 2 of 20 on the hard one.

GPT-6 Astra scores 99.9% or 62.7% on ARC-AGI-3, depending how you ask
OpenAI's GPT-6 Astra hit 99.9% on ARC-AGI-3 under OpenAI's own test setup, and 62.7% under ARC Prize's neutral one. The gap comes from what the model is allowed to remember.
The daily brief
Three to five stories a day, and what each one means for the people who build software. Free, no spam.