📊 Full opportunity report: The Rise Of Inkling: A New Era For Thinking Machines on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Thinking Machines has released Inkling, a 975-billion-parameter multimodal AI model, on Hugging Face. The model can process text, images, and audio within a large context window but requires substantial hardware. Its benchmarks and licensing details remain unverified.

Thinking Machines has released Inkling, a 975-billion-parameter multimodal model, on Hugging Face, offering new capabilities for processing text, images, and audio within a claimed one-million-token context window. This release marks a significant step in open AI models, though its hardware demands and lack of independent benchmarks limit immediate widespread use.

Inkling is described as a decoder-only Mixture-of-Experts model with 975 billion total parameters and 41 billion active during processing. It was trained on 45 trillion tokens across multiple data types, including text, images, audio, and video, though these figures have not been independently verified. The architecture employs 256 experts, with dynamic routing and attention mechanisms designed to optimize multimodal reasoning.

Hardware requirements are substantial: Hugging Face states that the BF16 checkpoint needs approximately 2 TB of VRAM, while the NVFP4 version requires about 600 GB, making it impractical for most consumer systems. The release includes support for popular inference frameworks, and developers can access the model via hosted services, but detailed licensing and performance benchmarks are not yet available. For more insights, see the original analysis.

At a glance
breakingWhen: announced July 2026
The developmentThinking Machines has launched Inkling, a large multimodal AI model, on Hugging Face, making it accessible for research and development despite high hardware requirements.
At a glance
announcementWhen: announced on Hugging Face; the source m…
The developmentThinking Machines has made its Inkling multimodal model available through Hugging Face with day-one support from several major inference frameworks.

Implications of Inkling’s Large-Scale Multimodal Capabilities

The release of Inkling signifies a major advancement in multimodal AI, potentially enabling more integrated and complex reasoning across different data types. Its open availability could accelerate research in scientific, media-analysis, and enterprise applications, though hardware constraints limit immediate adoption for many users. The lack of independent benchmarks and licensing clarity raises questions about its safety, reliability, and commercial viability in the near term.

Amazon

high VRAM graphics card for AI training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Large Multimodal AI Models and Recent Developments

Prior to Inkling, most large multimodal models have been proprietary or limited in scale, often requiring specialized hardware. The recent trend has been toward increasing model size and multimodal integration, but practical deployment has remained challenging due to hardware demands. The open release of Inkling by Thinking Machines builds on this trend, offering a model of unprecedented size and multimodal scope, though with limited independent evaluation or detailed licensing information.

“This model is huge.”

— Hugging Face spokesperson

Amazon

multimodal AI development hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Benchmark Results and Licensing Details

Independent evaluations of Inkling’s performance, safety, and robustness are not yet available. The licensing terms remain unspecified, and the impact of the large context window on real-world workloads is still unclear. How well the model performs on video or in consumer hardware settings also remains to be seen.

Amazon

large-scale AI model server

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Expected Developer Testing and Benchmarking Efforts

Early testing by researchers and developers will focus on latency, accuracy, and hardware compatibility, especially using the supported inference frameworks. Benchmark results, safety assessments, and licensing clarifications are anticipated in the coming months, which will determine the model’s practical viability and adoption scope.

Amazon

AI inference framework hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Inkling?

Inkling is a large, open multimodal AI model from Thinking Machines, capable of processing text, images, and audio, with 975 billion parameters and a large context window.

Can I run Inkling on my personal computer?

Unlikely. The model’s hardware requirements are extremely high, with estimates of 2 TB of VRAM for full deployment, making it suitable mainly for specialized hardware or hosted services.

Does Inkling support video processing?

The architecture accepts image inputs with a temporal dimension, but native video performance has not been evaluated or confirmed.

What are the licensing terms for Inkling?

The release describes Inkling as an open model but does not specify licensing details, usage restrictions, or whether the training data and code are publicly available.

When will independent benchmarks and safety evaluations be available?

There is no confirmed timeline, but early testing and evaluations by the developer community are expected to provide insights in the coming months.

Source: ThorstenMeyerAI.com

You May Also Like

At The Crossroads Of AI And Power: The China Open-Weight Doors

China considers restricting overseas access to advanced AI models amid US gating efforts, signaling a new AI power struggle with global implications.

World Model Readiness: Are You Ready for AI That Acts?

An emerging diagnostic tool evaluates organizations’ preparedness for AI systems that predict and act, marking a shift from language models to world models.

Best VPN for Streaming World Cup: Tested on July 1st.

On July 1st, cybersecurity firm Cybernews tested VPNs for streaming the World Cup, identifying the top performers for fans worldwide.

Memory Stopped Being a Commodity

Micron’s new long-term contracts signal a shift from memory being a flexible commodity to a pre-funded, strategic input for major buyers.