Learning to use local AI is exciting, overwhelming, and frustrating
The author explores running AI models locally on personal hardware, highlighting the privacy benefits over cloud services. Initial experiments with Hermes Agent on a Mac Studio reveal both potential and challenges.
Intelligence analysis by Gemini 2.5 Flash Lite

The article details the author's personal journey into using local AI agents, driven by privacy concerns and the desire for more control. It covers the setup of Hermes Agent on a powerful Mac Studio, the overwhelming choice of large language models, and early attempts at practical applications like generating daily briefings and organizing a Steam game library.
Imagine you have a super-smart helper robot that lives only in your computer, not on the internet. You can ask it to do chores like organizing your video games or telling you the weather. It's exciting because it's private, but sometimes it's tricky to figure out exactly what to ask it or how to set it up, like learning a new game.
Analysis
M5 Ultra Mac Studio
The author's exploration into local AI is anchored by the M5 Ultra Mac Studio, a machine boasting 256GB of unified memory. This substantial RAM capacity is presented as a key enabler for running large, parameter-heavy AI models locally, a significant departure from the constraints often faced with consumer-grade hardware. The article posits that Apple's new Mac desktops are being marketed with local AI capabilities in mind, suggesting a strategic push by the company to leverage its hardware for this burgeoning field. The sheer amount of memory available on the Mac Studio allows for experimentation with models that would be prohibitive or impossible on less capable systems, setting the stage for a more robust local AI experience.
This powerful hardware is crucial for the author's initial choice of the Qwen 3.8 Flash Next model, a 125-billion parameter model that consumes approximately 105GB of storage. The ability to run such a large model locally is a testament to the Mac Studio's capabilities and directly addresses the author's desire to explore the upper echelons of AI performance without the per-token costs associated with cloud-based services. The author also plans to test smaller Qwen models on other Apple devices and Windows machines, indicating a broader investigation into how hardware specifications influence the feasibility and performance of local AI.
Hermes Agent
At the core of the author's local AI setup is Hermes Agent, an open-source, self-hosted desktop application compatible with macOS, Windows, and Linux. The appeal of Hermes lies in its ability to function entirely with local LLMs, offering a completely free and private alternative to commercial cloud AI services. This self-hosted nature is central to the author's motivation, as it provides a "little gofer living solely in the box on my desk that answers only to me and not overlords from OpenAI, Google, Microsoft, or Anthropic." The agent's desktop interface is described as straightforward, including a model picker that simplifies the process of selecting and deploying various LLMs.
Hermes Agent's functionality extends to controlling AI models via a Telegram bot, allowing for remote interaction and task management. The author successfully configured Hermes to generate a daily morning briefing by scanning emails and calendars, a task that, while basic, demonstrates the agent's potential for automating routine information gathering. The initial setup challenges, such as ensuring the Mac was not asleep during scheduled tasks, highlight the practical considerations and troubleshooting involved in integrating local AI into daily workflows. The ability to customize and develop these automated tasks over time is presented as a key benefit of using a local, agentic system.
Qwen 3.8 Flash Next
The selection of the Qwen 3.8 Flash Next model represents a significant step in the author's exploration of local AI, driven by the ample resources of the Mac Studio. This model, with its 125 billion parameters and substantial 105GB size, is indicative of the high-end LLMs that can now be run locally, a capability that was largely out of reach for average users until recently. The author's decision to start with such a large model underscores a desire to test the limits of local AI performance and utility, unburdened by the token-based pricing structures of cloud providers.
The practical application of Qwen via Hermes Agent is illustrated through the task of reorganizing the author's Steam game library. This involved granting Hermes specific permissions, including the use of a Steam web API key, to access and categorize over 400 games. The agent's ability to quickly sort games by genre, while respecting user-defined custom categories, showcases the potential for AI to handle complex, data-intensive organizational tasks. The ease with which the API key could be revoked also points to the user's control over permissions and data access, a critical aspect of local AI's privacy advantage.
Key points
- Running AI models locally offers enhanced privacy compared to cloud-based services.
- The author is experimenting with local AI using Hermes Agent on a Mac Studio.
- Large language models require significant hardware resources, particularly RAM.
- Early use cases include generating daily briefings and organizing digital libraries.
- The process involves a learning curve and potential setup challenges.
The increasing accessibility of powerful local AI models on personal hardware could lead to a future where users have greater control over their data and AI interactions. This shift might foster more personalized and efficient workflows, as users can tailor AI agents to specific tasks without privacy concerns associated with cloud services.
The complexity and cost associated with setting up and running advanced local AI models, coupled with the steep learning curve for non-experts, could limit widespread adoption. Users might struggle with hardware requirements, software configuration, and identifying genuinely useful applications, leading to frustration and abandonment of the technology.


