Inkling-Small is a recently launched open-weights multimodal model from Thinking Machines Lab, described as an efficient alternative to its larger sibling, Inkling. Released on July 30, 2026, it is designed to deliver comparable performance to Inkling while using significantly less compute and offering lower latency.
Key Features
- Architecture & Scale: Inkling-Small is a Mixture-of-Experts (MoE) transformer model with 276 billion total parameters and 12 billion active parameters, compared to Inkling’s 975B total / 41B active scale.
- Multimodal Abilities: Like Inkling, the smaller model supports native multimodal understanding, handling input across text, images, and audio. Visual reasoning is enhanced via tools for image cropping, zooming, and programmatic inspection; audio input is represented via Mel-spectrograms.
- Variable Thinking Effort: Inkling-Small allows adjustment of reasoning “effort” at inference time, letting users trade off computation cost and performance depending on task complexity.
- Performance Benchmarks: Across various benchmarks (reasoning, agentic tasks, instruction following), Inkling-Small matches or exceeds Inkling in some areas—especially reasoning and agentic coding—though it lags in knowledge coverage and factuality.
- Safety & Epistemics: The model inherits safety mechanisms from Inkling, including internal and external tests, refusal on harmful or dual-use requests, and calibration efforts to express confidence properly.
Who is it for?
Inkling-Small is suited for business owners, enterprises, and development teams needing a model that balances capability and cost. Typical use cases include:
- Coding workflows, tool use automation, or internal agents where efficient reasoning output is crucial rather than maximum factual breadth.
- Multimodal content tasks—processing documents, charts, speech or audio content—where lightweight resource consumption is desirable.
- Scenarios where latency or infrastructure cost matter and the organization has some flexibility in acceptable model size.
Pricing
- Inkling-Small’s output token cost is estimated at $1.20 per million tokens, while Inkling’s full model costs approximately $4.05 per million tokens output. These pricing figures help highlight cost savings when using the smaller model.
- Both Inkling and Inkling-Small are available via Thinking Machines’ platform Tinker, including the “Tinker Playground” for audio, image, and text interaction.
- Full and fine-tuning weights are publicly released (e.g. on Hugging Face), making Inkling-Small accessible for organizations wanting to adapt or deploy the model themselves.
Final thoughts
Inkling-Small represents a clear step in Thinking Machines Lab’s strategy to provide efficient, open-weights models that are practical for real-world business use. Its lower active parameter count reduces resource requirements, while maintaining strong multimodal reasoning and safety features. However, trade-offs remain—especially in domains requiring deep factual knowledge or maximum knowledge coverage. For organizations where cost, latency, or infrastructure constraints are critical, Inkling-Small may offer a compelling balance. For tasks demanding state-of-the-art factual accuracy or broad knowledge, the larger Inkling or other specialized models might still be more appropriate.
Visit thinkingmachines.ai/news/inkling-small for more.
