Inkling is the latest open-weights multimodal foundation model from Thinking Machines Lab, released on July 15, 2026. Built as a mixture-of-experts transformer, Inkling features 975 billion parameters with 41 billion active, supports a context length of up to one million tokens, and accommodates inputs in text, image, and audio formats. A smaller variant, Inkling-Small, shares the same foundational architecture with 276 billion parameters and 12 billion active parameters, offering many of the same strengths at lower compute cost and latency.
Key Features
- Generalist Capabilities: Inkling excels across diverse domains—reasoning, coding, instruction-following, factual tasks, image and audio understanding—rather than optimizing for a single benchmark. Its design aims for broad adaptability rather than peak performance in any particular niche.
- Agentic Tool Use and Coding: The model demonstrates strong capabilities in writing and executing code, interacting with applications through embedded browser agents, and generating multi-page artifacts or apps from prompts. It also supports long, refinement-based workflows such as multiplayer game development guided by feedback loops.
- Controllable Thinking Effort: Users can adjust the trade-off between computation cost and performance through a parameter called “effort.” On many benchmarks, Inkling matches high-end open models while using significantly fewer active tokens.
- Multimodality: Inkling processes image, audio, and text inputs natively. It performs visual reasoning on charts, diagrams, and mathematical content, transcribes and understands speech, and integrates audio and image content in dynamic tasks. Tool-support allows transformations such as zooming and cropping within visual inputs.
- Epistemics and Safety: The model is trained for calibration (confidence estimation), robust instruction-following, and handling uncertain or censored content. Safety evaluations show Inkling excels in refusal benchmarks for harmful requests, and it was audited by external safety testers. The model is resistant to overconfidence and misuse in many scenarios.
Who is it for?
Inkling is designed for business owners, developers, and decision-makers who require a foundation model that can be customized to domain-specific workflows. It is particularly suited for organizations with needs in coding assistants, long-context document processing, multimodal analysis (e.g. image and audio tasks), AI agents or tools requiring web interaction, and applications where cost or latency are constraints. Inkling-Small is targeted at those who want similar core strengths with lower resource demands, ideal for experimentation, deployment in less powerful environments, or workloads sensitive to inference cost.
Pricing
Inkling is available via the Tinker platform, Thinking Machines Lab’s API for model customization, fine-tuning, and inference. For a limited time, users can access Inkling and Inkling-Small at 50% discount. Pricing is usage-based, billed per million tokens. For Inkling at a 64K context window, input (“prefill”) tokens cost approximately $1.87/M, with output (“sample”) tokens at $4.68/M and training (forward + backward passes) at $5.61/M. Extended 256K context incurs higher token costs roughly double those of the 64K context variant. Inkling-Small follows a similar structure but with lower per-token rates: $0.58 for prefills (64K context), $1.44 for sample-token output, and similar scaled train costs. Context window sizes of 64K and 256K tokens are supported. Token-cache mechanisms offer heavily discounted prefill rates when inputs hit cache. Storage of checkpoints is priced separately, typically at around $0.10 per gigabyte per month.
Final thoughts
Inkling represents a carefully balanced open-weights model that emphasizes versatility and accessibility alongside capability. It won’t be the absolute top performer in every benchmark compared to closed-source frontier models, but its strengths lie in broad support for multimodal tasks, adjustable computation trade-offs, and open distribution under an Apache 2.0 license. For businesses prioritizing customization, control, domain specificity, or cost containment, Inkling offers a viable and compelling option. Organizations considering adoption should assess whether they need the extended context variants, prepare for infrastructure demands if self-hosting, and plan for fine-tuning workflows to unlock Inkling’s full potential in their use cases.
Visit the official website for more.
Keep up to date with our stories on LinkedIn, Twitter, Facebook and Instagram.
