Mira Murati’s Thinking Machines ships Inkling: 975B open weights with Apache 2.0
A former OpenAI CTO’s first model targets frontier-class open access, plus quantized and fine-tunable deployment paths.

Mira Murati, former OpenAI CTO, founded Thinking Machines Lab in early 2025 and released its first frontier-class open weights model, code-named Inkling. Decision-makers now have a new, permissively licensed 975B American open-weights option that pushes competition against both proprietary labs and Chinese model houses.
Mira Murati, former OpenAI CTO, just put her name on a frontier open weights release: Thinking Machines Lab’s Inkling, a 975 billion parameter model released on Wednesday under a highly permissive Apache 2.0 license. That license is the headline for anyone who cares about control, portability, and downstream customization. Because Inkling is open in the weights sense and permissive in the terms sense, end users can fine tune it for specific use cases instead of being locked into one vendor’s API.
The scale is real, too. Running Inkling at native 16-bit precision needs more than two terabytes of GPU memory, which the company pegs as roughly available in around eight of Nvidia’s B300 accelerators or sixteen H200s. If you do not have that class of hardware, Thinking Machines also released a NVFP4 quantized version of the model intended to run on half the GPUs. The upshot: Murati is not just joining the open-weights party, she is showing up with a heavyweight.
In the open model market, options have indeed been limited outside of the Chinese model houses. Thinking Machines positions Inkling as the largest American open weights model to date, and it frames the model as comparable to Chinese systems like DeepSeek V4, GLM 5.2, and Kimi K2.6 in size and capabilities. The company also claims Inkling is competitive with those models across a variety of workloads, but its benchmark charts show it trailing proprietary systems like Anthropic’s Claude and OpenAI’s GPT. That matters for decision-makers because it clarifies the battlefield: Inkling is not necessarily aiming to dethrone every proprietary model, it is aiming to expand what teams can build on without treating API access like a toll road.
Thinking Machines describes Inkling as highly adaptable, designed for developers building AI apps while still usable for general purpose applications like chat bots. Under the hood, it is built around a mixture of experts (MoE) architecture, and the company says the MoE inspiration comes from DeepSeek-V3. But it also claims it trained Inkling from scratch using Nvidia GB300 NVL72 systems and 45 trillion tokens worth of text, images, audio, and video. The model includes 256 routed exports and two shared ones. It generates each token by six experts, totaling about 41 billion parameters. In practice, that design choice is the kind of thing that can make a model feel “big” on capability while keeping compute usage manageable enough to be commercially deployable, though the raw requirements and token costs still force hard budgeting.
A key product angle here is usability for builders, not just raw model bragging rights. Inkling ships with support for a million-token context, which the source describes as the model’s short-term memory. That is meant to help with tasks like wrangling large code bases and needle-in-the-haystack search problems. Thinking Machines also leans hard into a customization story through its Tinker platform. The company says Tinker offers tools for customization and fine tuning, and it even claims the model can write its own fine tuning scripts to refine behavior, teach itself new skills, and evaluate its abilities. Separate but related: Inkling is being released as a “reasoning model,” trained using reinforcement learning to use chain of thought before responding. The company says it tuned Inkling to use these thinking tokens more efficiently, and it claims Inkling matches Nvidia’s Nemotron 3 Ultra, up to now the largest and most capable American open weights model out there at 550 billion parameters, on Terminal Bench 2.1 using roughly a third the tokens.
That “thinking tokens” detail should put financial guardrails on anyone planning to deploy Inkling at scale. Those tokens are billed like any other, so longer deliberation can mean larger bills. Even if you own the model and can control your inference setup, the operational reality is the same: if the product uses more compute to reason, the cost comes due. And if you are evaluating vendors, that is the subtle distinction between “it can do it” and “it can do it profitably,” especially for apps where latency and cost are tied to user experience.
From an integration perspective, Inkling is available starting today on Thinking Machines’ Tinker platform, with the company also working to bring the model to third-party API services including TogetherAI, Fireworks, Modal, Databricks, and Baseten. For teams that prefer to run models themselves, Inkling is available for download on popular model repos like Hugging Face, and it claims support for a broad range of inference engines including vLLM, SGLang, Miles, TokenSpeed, and Llama.cpp. At launch, that “works with your stack” stance is not marketing fluff. It changes the adoption math because it reduces switching costs and lets teams prototype quickly without waiting for one closed ecosystem to bless them.
Finally, Murati’s first release comes with an apparent roadmap. Thinking Machines says Inkling is the first of several new models under development, and alongside its flagship model it is previewing Inkling-Small, a 276-billion-parameter MoE model with 12 billion active parameters for users prioritizing latency over throughput and quality. The company also says it is finalizing the model and plans to release its weights once testing is complete. For executives watching AI spending, this is a strategic signal: open weights is no longer just a curiosity for researchers. With permissive licensing, long context, and multiple deployment paths, it becomes a procurement and platform decision that can reshape how teams allocate compute budgets, negotiate API contracts, and design product architectures.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology

Nvidia and Wistron will build Blackwell AI servers in Texas, Nikkei Asia reports
A Texas manufacturing plan for Blackwell AI servers ties Nvidia's next platform rollout to Wistron's local capacity and supply chain risk.

Meta tests StoryKit bedtime stories in select regions to measure parent response
The experiment is regional, and the real question is how quickly parents adopt AI storytelling for kids.

Range Rover GT is not a Velar EV replacement, spy tests at Arctic Circle confirm
The EV plan is real, but the direction was misread for months. Here is the actual story.
