Felix Pinkston
Aug 10, 2026 11:02
Meta’s Muse Glimmer 30B, a 30B parameter open AI mannequin, now runs effectively on AMD {hardware}, enabling highly effective native AI functions.
Meta’s newest AI mannequin, Muse Glimmer 30B, is setting a brand new commonplace for native inference, and AMD {hardware} is positioned because the prime platform for operating it. This 30-billion-parameter dense mannequin, launched below the Apache 2.0 license, is optimized for agentic and coding workloads, giving builders a strong device for native AI functions with out counting on cloud infrastructure.
The mannequin is particularly designed to leverage AMD’s Ryzen™ AI Max+ processors and Radeon™ AI PRO R9700 graphics playing cards. Benchmarks point out spectacular efficiency: as much as 24 tokens per second on a Ryzen AI Max+ 395 processor and as much as 53 tokens per second on a single Radeon AI PRO R9700 GPU. These figures include dFlash enabled and have been examined utilizing open frameworks like llama.cpp, additional highlighting the flexibility of the system.
Muse Glimmer’s open-weight structure helps native workflows, a important benefit for privacy-sensitive functions. In contrast to cloud-first AI options that may introduce latency and safety dangers, operating the mannequin domestically retains information and operations below direct management. This aligns with the rising development of decentralized AI adoption, the place builders prioritize on-device options for sooner, safer, and extra cost-efficient functions. Early reviews from the Unsloth neighborhood counsel the mannequin requires about 18GB of RAM for inference, making it accessible for high-end shopper PCs and workstations.
Why Muse Glimmer 30B Issues
The 30B mannequin is a significant leap for agentic AI. It’s designed to deal with advanced, multi-step workflows that demand context retention, device utilization, and adaptive decision-making. With help for 128K+ context, Muse Glimmer might energy functions like coding assistants, workflow automation, and personal AI instruments. Its coaching additionally emphasizes security, with mechanisms to withstand oversharing and immediate injection assaults.
Builders can get began with Meta’s LM Studio for simple mannequin deployment on AMD {hardware}. Techniques with 32GB+ Variable Graphics Reminiscence (VGM) or VRAM are really useful, permitting customers to run the mannequin in minutes. For integration into present functions, the Lemonade platform presents a light-weight API that simplifies deployment whereas maintaining information native. This embeddable binary strategy opens the door to quite a lot of real-world use circumstances, from analysis instruments to offline assistants.
Open-Supply and Business Flexibility
Muse Glimmer’s Apache 2.0 licensing provides builders in depth freedom. They will modify, redistribute, and combine the mannequin into industrial functions with out restrictive phrases. Group contributors have already launched GGUF quantized variations on Hugging Face, enabling extra environment friendly inference on a variety of {hardware}.
This launch positions Meta as a frontrunner within the push for native AI options. By specializing in agentic workloads and prioritizing privateness, Muse Glimmer 30B fills a important hole out there. It’s particularly related as companies and builders look to transition from cloud-dependent AI to on-device options that supply larger management and decrease working prices.
What’s Subsequent?
Muse Glimmer 30B is a part of AMD’s imaginative and prescient for the “Agentic PC,” a platform that strikes past remoted AI options to host persistent intelligence. With additional software program optimizations and ecosystem growth anticipated, the efficiency of Muse Glimmer on AMD {hardware} will doubtless enhance additional. For builders and enterprises searching for to discover the following section of native AI, this mannequin and its {hardware} companions symbolize a compelling place to begin.
Picture supply: Shutterstock

