Thinking Machines Lab is releasing Inkling-Small, an efficient, open-weight model designed to deliver performance comparable to Inkling while being one-quarter of its size. The Mixture-of-Experts transformer contains 276 billion total parameters, with 12 billion active. Its full weights are being released on Hugging Face, with fine-tuning offered through Tinker and text, image, and audio chat available in Tinker Playground.
The model combines a context window of up to one million tokens with variable thinking effort from minimal to xhigh, allowing developers to trade compute for performance. It was trained on NVIDIA GB300 NVL72 systems and uses substantially less compute than Inkling, according to Thinking Machines. That lower requirement positions it for experimentation, coding, tool use, fine-tuning, and real applications.
Today, we are releasing Inkling-Small.
— Thinking Machines (@thinkymachines) July 30, 2026
Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available.https://t.co/NzFVYVkuQI
Fine-tune it on Tinker today, or chat with…
Thinking Machines began Inkling-Small after training Inkling, giving the lab room to revise its pre-training data mix and machine-learning recipe. An earlier preview checkpoint was post-trained partly through on-policy distillation with Inkling as the teacher, followed by two weeks of scaled agentic coding reinforcement learning. The lab says the resulting model overtakes Inkling on reasoning and agentic coding benchmarks, although Inkling remains stronger in knowledge coverage and factuality.
Reported results include 31.6 percent on the text-only Humanity's Last Exam, compared with 29.7 percent for Inkling. Inkling-Small also reached 80.2 percent on SWEBench Verified, 55.9 percent on the public SWEBench Pro evaluation, and 64.7 percent on Terminal-Bench 2.1 with its best harness. Thinking Machines ran its evaluations at an effort of 0.99 and a temperature of 1.0, with disclosed caveats regarding internal harness and formatting.
Inkling-Small keeps Inkling's natively multimodal, encoder-free design. Audio is converted to dMel spectrograms, while images are split into 40-by-40-pixel patches and transformed by a four-layer hMLP before being processed with text tokens. The model can use Python for visual work, combining reasoning with cropping, zooming, and programmatic inspection of dense documents and charts.
Safety post-training follows Inkling's recipe and is backed by internal testing and red-teaming by trusted external partners. Thinking Machines reports 98.4 percent on StrongREJECT, a 71.6 percent adversarial refusal rate on FORTRESS, and a 96.9 percent benign answer rate. Inkling and Inkling-Small are currently offered on Tinker with a limited-time discount.