AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: The Rise Of Inkling: A New Era For Thinking Machines on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

TL;DR

Thinking Machines has released Inkling, a 975-billion-parameter multimodal AI model, on Hugging Face. The model can process text, images, and audio within a large context window but requires substantial hardware. Its benchmarks and licensing details remain unverified.

Thinking Machines has released Inkling, a 975-billion-parameter multimodal model, on Hugging Face, offering new capabilities for processing text, images, and audio within a claimed one-million-token context window. This release marks a significant step in open AI models, though its hardware demands and lack of independent benchmarks limit immediate widespread use.

Inkling is described as a decoder-only Mixture-of-Experts model with 975 billion total parameters and 41 billion active during processing. It was trained on 45 trillion tokens across multiple data types, including text, images, audio, and video, though these figures have not been independently verified. The architecture employs 256 experts, with dynamic routing and attention mechanisms designed to optimize multimodal reasoning.

Hardware requirements are substantial: Hugging Face states that the BF16 checkpoint needs approximately 2 TB of VRAM, while the NVFP4 version requires about 600 GB, making it impractical for most consumer systems. The release includes support for popular inference frameworks, and developers can access the model via hosted services, but detailed licensing and performance benchmarks are not yet available. For more insights, see the original analysis.

At a glance
breakingWhen: announced July 2026
The developmentThinking Machines has launched Inkling, a large multimodal AI model, on Hugging Face, making it accessible for research and development despite high hardware requirements.
At a glance
announcementWhen: announced on Hugging Face; the source m…
The developmentThinking Machines has made its Inkling multimodal model available through Hugging Face with day-one support from several major inference frameworks.

Implications of Inkling’s Large-Scale Multimodal Capabilities

The release of Inkling signifies a major advancement in multimodal AI, potentially enabling more integrated and complex reasoning across different data types. Its open availability could accelerate research in scientific, media-analysis, and enterprise applications, though hardware constraints limit immediate adoption for many users. The lack of independent benchmarks and licensing clarity raises questions about its safety, reliability, and commercial viability in the near term.

GIGABYTE Radeon™ AI PRO R9700 AI TOP 32G Graphics Card, Turbo Fan Cooling System, 32GB GDDR6, GV-R9700AI TOP-32GD Video Card

GIGABYTE Radeon™ AI PRO R9700 AI TOP 32G Graphics Card, Turbo Fan Cooling System, 32GB GDDR6, GV-R9700AI TOP-32GD Video Card

  • Powered by Radeon AI PRO R9700: Enhanced RDNA 4 architecture and AI accelerators
  • 32GB GDDR6 Memory: Supports large, complex projects
  • PCIe Gen 5 Support: Fast data transfer speeds

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Large Multimodal AI Models and Recent Developments

Prior to Inkling, most large multimodal models have been proprietary or limited in scale, often requiring specialized hardware. The recent trend has been toward increasing model size and multimodal integration, but practical deployment has remained challenging due to hardware demands. The open release of Inkling by Thinking Machines builds on this trend, offering a model of unprecedented size and multimodal scope, though with limited independent evaluation or detailed licensing information.

“This model is huge.”

— Hugging Face spokesperson

Yahboom Raspberry Pi 5 ROS2 Robot Car 360°Movement, AI Vision & Tracking, Integrated Multimodal Large AI Model OpenRouter, AI Voice Interaction (Superior Without RPi5)

Yahboom Raspberry Pi 5 ROS2 Robot Car 360°Movement, AI Vision & Tracking, Integrated Multimodal Large AI Model OpenRouter, AI Voice Interaction (Superior Without RPi5)

  • Powerful Raspberry Pi 5 Control: Enhanced processing, multimedia, and AI performance
  • Large AI Model Integration: Advanced human-computer interaction and environmental perception
  • Multiple Control Options: App, PC, remote control, and FPV transmission

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Benchmark Results and Licensing Details

Independent evaluations of Inkling’s performance, safety, and robustness are not yet available. The licensing terms remain unspecified, and the impact of the large context window on real-world workloads is still unclear. How well the model performs on video or in consumer hardware settings also remains to be seen.

NIMO AI NAS, Agentic Computer Mini PC and AI Server, AMD Ryzen 7 PRO 8845HS

NIMO AI NAS, Agentic Computer Mini PC and AI Server, AMD Ryzen 7 PRO 8845HS

  • AI & LLM Local Deployment: Powered by AMD Ryzen 7 PRO 8845HS and RTX 5060 GPU
  • 4K/8K Video Editing Hub: Supports real-time high-resolution video editing
  • Virtualization & Docker Support: Handles multiple VMs and Docker containers smoothly

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Expected Developer Testing and Benchmarking Efforts

Early testing by researchers and developers will focus on latency, accuracy, and hardware compatibility, especially using the supported inference frameworks. Benchmark results, safety assessments, and licensing clarifications are anticipated in the coming months, which will determine the model’s practical viability and adoption scope.

AI in Embedded Systems: Types, Techniques, Machine Learning, Model Training vs. On-device Inference, Algorithms, Frameworks and Tools.

AI in Embedded Systems: Types, Techniques, Machine Learning, Model Training vs. On-device Inference, Algorithms, Frameworks and Tools.

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Inkling?

Inkling is a large, open multimodal AI model from Thinking Machines, capable of processing text, images, and audio, with 975 billion parameters and a large context window.

Can I run Inkling on my personal computer?

Unlikely. The model’s hardware requirements are extremely high, with estimates of 2 TB of VRAM for full deployment, making it suitable mainly for specialized hardware or hosted services.

Does Inkling support video processing?

The architecture accepts image inputs with a temporal dimension, but native video performance has not been evaluated or confirmed.

What are the licensing terms for Inkling?

The release describes Inkling as an open model but does not specify licensing details, usage restrictions, or whether the training data and code are publicly available.

When will independent benchmarks and safety evaluations be available?

There is no confirmed timeline, but early testing and evaluations by the developer community are expected to provide insights in the coming months.

Source: ThorstenMeyerAI.com

BABY SHOWER & RE

Baby shower & registry season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Exploring Kong Tao’s Shift From ByteDance To Robotics Work With Lei Jun

Tsinghua PhD Kong Tao reportedly departs ByteDance to work with Lei Jun in robotics, though key details remain unconfirmed.

Windows 11’S Built-in Weather App Wastes More Than 1 GB Of RAM

Research shows Windows 11’s built-in Weather app consumes more than 1 GB of RAM, raising concerns about efficiency and system performance.

The Future Of AI: Implementing Nunchaku 4-Bit Diffusion In Diffusers

Hugging Face has added native support for Nunchaku Lite 4-bit checkpoints in Diffusers, reducing memory use and increasing inference speed without extra engines.

Unveiling Grok 4.6: SpaceXAI’s AI Model Designed For Long-Running, Knowledge-Heavy Tasks

SpaceXAI announced Grok 4.6, a new AI model with a 500K context window, designed for long-running tasks like coding and knowledge work. Details on access and performance remain pending.