Inkling is the latest open-weights multimodal foundation model from Thinking Machines Lab, released on July 15, 2026. Built as a mixture-of-experts transformer, Inkling features 975 billion parameters with 41 billion active, supports a context length of up to one million tokens, and accommodates inputs in text, image, and audio formats. A smaller variant, Inkling-Small, shares the same foundational architecture with 276 billion parameters and 12 billion active parameters, offering many of the same strengths at lower compute cost and latency.
Inkling is designed for business owners, developers, and decision-makers who require a foundation model that can be customized to domain-specific workflows. It is particularly suited for organizations with needs in coding assistants, long-context document processing, multimodal analysis (e.g. image and audio tasks), AI agents or tools requiring web interaction, and applications where cost or latency are constraints. Inkling-Small is targeted at those who want similar core strengths with lower resource demands, ideal for experimentation, deployment in less powerful environments, or workloads sensitive to inference cost.
Inkling is available via the Tinker platform, Thinking Machines Lab’s API for model customization, fine-tuning, and inference. For a limited time, users can access Inkling and Inkling-Small at 50% discount. Pricing is usage-based, billed per million tokens. For Inkling at a 64K context window, input (“prefill”) tokens cost approximately $1.87/M, with output (“sample”) tokens at $4.68/M and training (forward + backward passes) at $5.61/M. Extended 256K context incurs higher token costs roughly double those of the 64K context variant. Inkling-Small follows a similar structure but with lower per-token rates: $0.58 for prefills (64K context), $1.44 for sample-token output, and similar scaled train costs. Context window sizes of 64K and 256K tokens are supported. Token-cache mechanisms offer heavily discounted prefill rates when inputs hit cache. Storage of checkpoints is priced separately, typically at around $0.10 per gigabyte per month.
Inkling represents a carefully balanced open-weights model that emphasizes versatility and accessibility alongside capability. It won’t be the absolute top performer in every benchmark compared to closed-source frontier models, but its strengths lie in broad support for multimodal tasks, adjustable computation trade-offs, and open distribution under an Apache 2.0 license. For businesses prioritizing customization, control, domain specificity, or cost containment, Inkling offers a viable and compelling option. Organizations considering adoption should assess whether they need the extended context variants, prepare for infrastructure demands if self-hosting, and plan for fine-tuning workflows to unlock Inkling’s full potential in their use cases.
Keep up to date with our stories on LinkedIn, Twitter, Facebook and Instagram.
Built by our team member Maziar Foroudian, Mazi is an intelligent agent designed to research across trusted websites and craft insightful, up-to-date content tailored for business professionals.
Learn why choosing an MAS-licensed cryptocurrency exchange ensures maximum security and transparency.
How telematics and fleet technology help businesses cut costs, reduce downtime, improve productivity, and make smarter operational decisions.
Essential guide for Australian SMEs on navigating the AI search revolution and maintaining online visibility in 2026.
Planet Ark’s revamped recycling equipment catalogue helps businesses reduce waste, lower costs, improve recycling performance, and achieve sustainability goals efficiently.
How a maths-based judging system is eliminating bias and reshaping trust in global business awards.