Skip to content

AstaBrief 8B: Ai2's open model writes cited reports

Ai2 released AstaBrief 8B, an Apache 2.0 model on Qwen3-8B that writes cited research reports in 51.1 seconds, 3.5x faster than Asta's Thinking mode.

By Tech AI Wire Team

3 min read

XLinkedIn
Screenshot of the allenai/AstaBrief_8B model page on Hugging Face, showing its Apache 2.0 license tag and the start of its model card.
Average time per Asta report
Fast mode (AstaBrief 8B)
51.1s
Thinking mode (Claude-powered)
178.5s

The Allen Institute for AI (Ai2) released AstaBrief 8B on October 2, 2026, a small open model that turns research papers into a cited report. It writes a report in 51.1 seconds on average, about 3.5 times faster than the Claude-powered mode it sits beside. The weights are under the Apache 2.0 license, so a team can run it on its own hardware and ship it in a product.

AstaBrief is now the "Fast mode" in Asta, Ai2's platform for answering research questions from the scientific literature. The slower "Thinking mode," which runs on Anthropic's Claude, stays available.

What AstaBrief does

AstaBrief takes two inputs: a research question and excerpts from papers that a search step has already found. It returns a written report with citations to those excerpts, according to the Ai2 announcement.

The speed comes from doing the job in one pass. Older report pipelines summarize sources, cluster them and then draft section by section. AlphaSignal reports that AstaBrief skips those intermediate stages and writes the whole report at once.

The model has 8 billion parameters, the adjustable values a model learns in training. It is built on Qwen3-8B, an open model from Alibaba's Qwen team. Its Hugging Face model card lists a context length of 16,000 tokens. A token is a word or part of a word, so that is roughly the reading budget for the question and the paper excerpts together.

How Ai2 trained it

Ai2 started from about 90,000 real research queries from Asta users, filtered for quality. From those it built 47,000 examples for supervised fine-tuning, where the model learns by copying good answers. It then ran about 6,000 preference pairs through direct preference optimization (DPO). In DPO the model sees two reports for the same question and learns to favor the better one.

The example reports came from other models: Claude 3.5 and 3.7 Sonnet, o3, o4-mini, GPT-4.1, DeepSeek-V3 and DeepSeek-R1. GPT-4.1 and DeepSeek-R1 acted as judges to pick the better report in each pair. The model card says those judges agreed with human preferences 95% of the time.

One lesson stood out. Filtering the training data for reports dense with citations was "the single biggest driver of grounding quality," Ai2 says, as quoted by AlphaSignal. Grounding here means the report's claims trace back to the papers it cites.

Training ran on eight Nvidia H100 GPUs. The model card says inference fits on a single high-memory consumer GPU and works with the Transformers and vLLM libraries.

How it scores

Ai2 tested AstaBrief on SQABench-CS2, a set of 200 computer-science research questions. AlphaSignal calls the same set ScholarQA-CS2.

Metric on SQABench-CS2AstaBrief 8B
Average score87.0
Citation precision90.5
Ingredient recall (expected key points covered)90.2
Answer precision89.0
Citation recall78.2

Citation recall is the weak spot. It measures how many of the claims that need a citation actually get one. On a second test, DeepScholarBench, built from 63 arXiv-based queries, AstaBrief scored 53.5, according to AlphaSignal.

What it cannot do

AstaBrief only writes. It does not search the open web, break a question into sub-questions or run follow-up searches, AlphaSignal reports. Something else has to find the papers first. AlphaSignal also warns that its citations may overstate what a paper actually found.

The comparisons are dated, too. Unite.AI notes that training and evaluation were done in 2025 and were not rerun against current frontier models.

What this means for developers

The useful part is the full release. Ai2 published the weights, the training dataset (AstaBrief_DPO_Mix), code examples and hyperparameters. That makes AstaBrief a recipe as much as a model. A team building a report writer for legal documents, support tickets or internal wikis can copy the setup with its own data.

If you already run a retrieval pipeline, AstaBrief fits at the end of it. Your search step finds passages, and AstaBrief writes the cited summary. Because the weights are open, the documents never have to leave your servers. For many companies that is the main reason to pick an 8B model over a hosted one.

Test the weak spot before you ship it. A citation recall of 78.2 means some claims in a report will arrive without a source. Sample reports by hand, check every citation against the passage it points to, and keep a person in the loop for anything a reader will act on.

Sources

  1. Open-sourcing AstaBrief, the fast report-generation model in Asta - Ai2
  2. allenai/AstaBrief_8B - Hugging Face
  3. Ai2 Open-Sources AstaBrief 8B for Fast Scientific Report Generation - Unite.AI
  4. Ai2's AstaBrief 8B Writes Cited Research Reports 3.5x Faster in One Pass - AlphaSignal

Related articles

Screenshot of the Qwen Research License Agreement file for Qwen-Image-2.1 on Hugging Face, showing its September 20, 2026 release date, the definition of non-commercial use and the clause barring commercial use.
AI & LLMs

Qwen's research license, explained

The Qwen Research License lets anyone download Qwen-Image-2.1 but limits it to research or evaluation. Commercial use needs a separate license from Qwen.

The daily brief

Three to five stories a day, and what each one means for the people who build software. Free, no spam.