AstaBrief 8B: Ai2's open model writes cited reports
Ai2 released AstaBrief 8B, an Apache 2.0 model on Qwen3-8B that writes cited research reports in 51.1 seconds, 3.5x faster than Asta's Thinking mode.
3 min read

- Fast mode (AstaBrief 8B)
- 51.1s
- Thinking mode (Claude-powered)
- 178.5s
The Allen Institute for AI (Ai2) released AstaBrief 8B on October 2, 2026, a small open model that turns research papers into a cited report. It writes a report in 51.1 seconds on average, about 3.5 times faster than the Claude-powered mode it sits beside. The weights are under the Apache 2.0 license, so a team can run it on its own hardware and ship it in a product.
AstaBrief is now the "Fast mode" in Asta, Ai2's platform for answering research questions from the scientific literature. The slower "Thinking mode," which runs on Anthropic's Claude, stays available.
What AstaBrief does
AstaBrief takes two inputs: a research question and excerpts from papers that a search step has already found. It returns a written report with citations to those excerpts, according to the Ai2 announcement.
The speed comes from doing the job in one pass. Older report pipelines summarize sources, cluster them and then draft section by section. AlphaSignal reports that AstaBrief skips those intermediate stages and writes the whole report at once.
The model has 8 billion parameters, the adjustable values a model learns in training. It is built on Qwen3-8B, an open model from Alibaba's Qwen team. Its Hugging Face model card lists a context length of 16,000 tokens. A token is a word or part of a word, so that is roughly the reading budget for the question and the paper excerpts together.
How Ai2 trained it
Ai2 started from about 90,000 real research queries from Asta users, filtered for quality. From those it built 47,000 examples for supervised fine-tuning, where the model learns by copying good answers. It then ran about 6,000 preference pairs through direct preference optimization (DPO). In DPO the model sees two reports for the same question and learns to favor the better one.
The example reports came from other models: Claude 3.5 and 3.7 Sonnet, o3, o4-mini, GPT-4.1, DeepSeek-V3 and DeepSeek-R1. GPT-4.1 and DeepSeek-R1 acted as judges to pick the better report in each pair. The model card says those judges agreed with human preferences 95% of the time.
One lesson stood out. Filtering the training data for reports dense with citations was "the single biggest driver of grounding quality," Ai2 says, as quoted by AlphaSignal. Grounding here means the report's claims trace back to the papers it cites.
Training ran on eight Nvidia H100 GPUs. The model card says inference fits on a single high-memory consumer GPU and works with the Transformers and vLLM libraries.
How it scores
Ai2 tested AstaBrief on SQABench-CS2, a set of 200 computer-science research questions. AlphaSignal calls the same set ScholarQA-CS2.
| Metric on SQABench-CS2 | AstaBrief 8B |
|---|---|
| Average score | 87.0 |
| Citation precision | 90.5 |
| Ingredient recall (expected key points covered) | 90.2 |
| Answer precision | 89.0 |
| Citation recall | 78.2 |
Citation recall is the weak spot. It measures how many of the claims that need a citation actually get one. On a second test, DeepScholarBench, built from 63 arXiv-based queries, AstaBrief scored 53.5, according to AlphaSignal.
What it cannot do
AstaBrief only writes. It does not search the open web, break a question into sub-questions or run follow-up searches, AlphaSignal reports. Something else has to find the papers first. AlphaSignal also warns that its citations may overstate what a paper actually found.
The comparisons are dated, too. Unite.AI notes that training and evaluation were done in 2025 and were not rerun against current frontier models.
What this means for developers
The useful part is the full release. Ai2 published the weights, the training dataset (AstaBrief_DPO_Mix), code examples and hyperparameters. That makes AstaBrief a recipe as much as a model. A team building a report writer for legal documents, support tickets or internal wikis can copy the setup with its own data.
If you already run a retrieval pipeline, AstaBrief fits at the end of it. Your search step finds passages, and AstaBrief writes the cited summary. Because the weights are open, the documents never have to leave your servers. For many companies that is the main reason to pick an 8B model over a hosted one.
Test the weak spot before you ship it. A citation recall of 78.2 means some claims in a report will arrive without a source. Sample reports by hand, check every citation against the passage it points to, and keep a person in the loop for anything a reader will act on.
Sources
Related articles

Qwen's research license, explained
The Qwen Research License lets anyone download Qwen-Image-2.1 but limits it to research or evaluation. Commercial use needs a separate license from Qwen.

Xiaomi MiMo-V2.6-Pro ships 1T open weights under MIT
Xiaomi's MiMo-V2.6-Pro has 1.02 trillion parameters, 42 billion active, a 1M-token context and MIT-licensed weights, but needs about 680 GB of GPU memory.

NVIDIA Kumo Tabular tops TabArena with open weights
NVIDIA's Kumo Tabular predicts from a labeled table in one pass, with no training run. It ranks first on TabArena at 1,950 Elo, in sizes from 28M to 215M.
The daily brief
Three to five stories a day, and what each one means for the people who build software. Free, no spam.