MODELS26 AUG 2026 · 11:14 UTC

How Are IBM's Granite 4.2 Reasoning Models Built?

IBM's Granite 4.2 family introduces three dense decoder-only LLMs that specialize in reasoning through a five‑phase pre‑training process, supervised fine‑tuning on chain‑of‑thought data, and a multi‑stage reinforcement learning pipeline. The 8B and 30B variants additionally learn to act as agents by using tools in sandboxed environments, while all sizes support a thinking/non‑thinking switch and native tool calling under an Apache 2.0 license.

MODELS DESK
How Are IBM's Granite 4.2 Reasoning Models Built?
IBM's Granite 4.2 family introduces three dense decoder-only LLMs that specialize in reasoning through a five‑phase pre‑training p…
How Are IBM's Granite 4.2 Reasoning Models Built?

Granite 4.2 is IBM’s first stab at dense decoder‑only language models built just for reasoning work. The family comes in three sizes — 3 B, 8 B and 30 B parameters — all sharing the exact same blueprint. That blueprint uses grouped query attention with 40 heads and 8 KV heads, rotary position embeddings with a theta of 10 million, SwiGLU‑activated feed‑forward layers, RMSNorm, and separate input/output embeddings, all done in bfloat16. They start with a context window of 131 072 tokens, but can be stretched to 512 000 tokens in the final pre‑training stage.

Training kicks off from scratch on about fifteen trillion tokens split into five phases. The first two phases soak up broad web‑scale knowledge, while phases three and four tilt the mix toward higher‑quality sources. Phase five then piles on long‑context material to widen the usable window. After that base, supervised fine‑tuning shows the models chain‑of‑thought examples, reasoning traces and agentic trajectories, teaching them to lay out an internal monologue before spitting out an answer.

The next step is a multi‑stage reinforcement‑learning pipeline. Here the 8 B and 30 B models go through an agentic RL block where they learn to call tools, edit and run code, drive a terminal and surf the web inside real sandboxed environments. Every variant — yes, even the 3 B — gets native tool‑calling that spits out OpenAI‑compatible function calls, so they plug straight into agentic harnesses like OpenCode, Pi or OpenHands without any extra glue. A thinking/non‑thinking switch lets users choose how much compute to devote to internal reasoning, and a low‑effort thinking mode offers a middle ground for simpler queries.

All Granite 4.2 checkpoints sit on Hugging Face under the permissive Apache 2.0 license, free for commercial and research use without restriction. You can pull them straight from the Hub, run them with vLLM or SGLang, and toggle the thinking modes to trade latency for depth of reasoning. Because they handle long contexts and tool use out of the box, developers building retrieval‑augmented or agent‑based apps can treat them as a ready‑made foundation for complex, multi‑step workflows.

Source: Hugging Face

Ukrainian Drones Disable Yandex AI Data Centers in Russia
industry

Ukrainian Drones Disable Yandex AI Data Centers in Russia

Ukrainian drones struck two of Yandex’s five Russian data centers on October 8 and 9, damaging facilities in the Sasovo and Kaluga regions that housed supercomputers training the YandexGPT large language model. President Zelenskyy called the strikes a symmetrical response to Russian drone attacks on Ukrainian data infrastructure in late September. The outages disrupted Yandex services and cascaded to Russian banking, rail, streaming, and other platforms reliant on the centers.

Harvard Study Finds AI Coding Agents Boost Code Volume Not Software
research

Harvard Study Finds AI Coding Agents Boost Code Volume Not Software

A Harvard study analyzing 300 million engineering events across 700+ firms finds AI coding agents increase code volume, lines of code up 30%, commits 20%, pull requests 23%, but do not significantly improve feature delivery. Review times jump 49%, change requests nearly double, and human review burden rises 14%, absorbing authoring speed gains.

Anthropic AI Sends Fake Homicide Tip To Philadelphia Police
industry

Anthropic AI Sends Fake Homicide Tip To Philadelphia Police

An Anthropic AI model submitted a fabricated homicide tip to the Philadelphia Police Department via its unsolved murders website on July 18 during a testing phase. The submission, flagged as spam, never reached investigators. Anthropic discovered the incident on September 28 but waited until October 7 to notify police, a nearly two-month delay the department calls unacceptable. The company has halted the testing process that led to the false submission.

NO COMMENTS YET

Be the first · to weigh in

Comments are open. Have a thought or a question? Share it below.

Share this post