Hermes Agent costs nothing and is one of the more expensive pieces of software you can run, and both of those statements are true at the same time. The agent itself is MIT-licensed; Nous Research's own FAQ says you "pay only for the LLM API usage from your chosen provider", and that local models are free to run. The bill comes from somewhere else: an agent does not send one message and wait for one answer. It plans, calls a tool, reads the result, calls another, and each of those steps is a separate request to a model that charges by the token.

The cost guides published this month list per-token prices and a VPS and stop there. What they miss is the shape of an agent's spending, which is what decides whether your month costs three dollars or a hundred. This piece is arithmetic on the official price lists and on one community measurement of what Hermes actually sends per request; it is not a lab test, and where a number is an estimate it says so.

Free software, metered thinking

There are three ways to feed a model to Hermes, and they have different cost shapes. The first is to bring your own API key. The providers documentation lists direct support for MiniMax, Anthropic, OpenAI, DeepSeek, Gemini and xAI, plus OpenRouter, which fronts a few hundred models behind one key, and "any service with an OpenAI-compatible API". You pay each provider's list price. Switching is a single hermes model command, which matters more than it sounds, because it means the price list is a menu rather than a decision.

The second route is Nous Portal, the company's own subscription. The free tier gives you free models only. Plus is $20 a month for $22 of credits, Super is $100 for $110, Ultra is $200 for $220, with rollover caps of $10, $50 and $100 respectively and access to 200-plus models and hosted tools. It is, in effect, a prepaid card with a 10 per cent bonus and a rate-limit upgrade; the underlying tokens still cost what they cost.

The third route is local. Point Hermes at Ollama on localhost:11434 and the tokens are free; the docs note that Ollama needs a context window of at least 64,000 tokens for Hermes to work properly. Ollama's own Hermes page suggests Gemma 4 for reasoning and code at around 16 GB of video memory and Qwen 3.6 at around 24 GB. That is a serious graphics card or a well-specified Mac, and the electricity is real: one published estimate puts a 450-watt card running around the clock at about $39 a month at average US rates, and a Mac at under $3. If you have followed our LocalAI setup on Windows or the Claude Code with Ollama walk-through, you already have most of this route built.

There is a fourth, half-route: managed hosting. Hostinger sells a pre-configured Hermes box at a promotional $5.99 a month, renewing at $11.99, with "pre-configured AI credits", web search and Telegram pairing included, paid upfront. It removes the API-key step, which for a non-technical user is most of the friction, at the cost of a credit balance you did not choose the size of. Our tips for non-technical Hermes users covers what that trade looks like day to day.

The number that matters is calls, not words

The useful measurement comes from a Hermes user who instrumented the agent and posted the results as a GitHub issue. Every request Hermes makes carries a fixed preamble of about 13,900 tokens: roughly 5,200 tokens of system prompt and 8,800 tokens describing the 31 tools the model may call. Add the conversation itself and a typical request ran to 17,000 to 23,000 tokens. Over one evening, three sessions made 207 API calls and sent about 3.9 million input tokens. The measurement was taken on an older release, v0.6.0, and the current September builds have changed the prompt, so treat the figures as the order of magnitude rather than the exact overhead.

Two things follow. The first is that your prompt is not the cost. Whether you type twelve words or two hundred, the request that leaves the machine is dominated by the same preamble, repeated every time the agent thinks. The second is that the price that matters is not the headline input price but the cached-input price, because that preamble is identical from one call to the next and every major provider now bills a repeated prefix at a steep discount, provided the system prompt stays stable between calls. Nous's own tips page says exactly this: keep the context files and memory the same and "subsequent messages in a session get cache hits that are significantly cheaper".

Here is that measured evening priced at each provider's current list, first as if nothing were cached and then as if the whole thing were. Reality sits between the two columns, closer to the second on a well-behaved session.

ModelInput / 1MCached input / 1MOutput / 1M3.9M input tokens, uncachedSame, fully cached
MiniMax M3 (MiniMax platform)$0.30$0.06$1.20$1.17$0.23
MiniMax M3 (via OpenRouter)$0.23$0.05$0.96$0.90$0.20
DeepSeek V4 Flash (via OpenRouter)$0.04not listed$0.10$0.16—
GPT-5.6 Luna$0.20$0.02$1.20$0.78$0.08
Gemini 3.5 Flash-Lite$0.30$0.03$2.50$1.17$0.12
Claude Haiku 4.5$1.00$0.10$5.00$3.90$0.39
Claude Sonnet 5$2.00$0.20$10.00$7.80$0.78
Claude Opus 5$5.00$0.50$25.00$19.50$1.95

Prices are from the MiniMax, Anthropic, OpenAI and Google pricing pages and the OpenRouter listings on 17 September 2026. Output tokens are extra and smaller: if each of those 207 calls produced 500 tokens of text, that is about 100,000 output tokens, which adds 12 cents on M3 and $2.50 on Opus. The MiniMax page still shows M3's standard price as a "permanent 50% off" the launch rate, and a higher tier applies to requests over 512,000 tokens of input, which an agent session can reach if you never compress it.

Multiply by thirty and the month comes into view. An evening like that every day costs about $35 on MiniMax M3 with no caching and about $7 with full caching; on Claude Opus 5 it is $585 and $59. That spread is the entire answer to "how much does Hermes cost": the model choice is a sixteen-times multiplier between M3 and Opus, and the cache is a five- to ten-times divider, and the agent sits in the middle sending the same preamble a couple of hundred times a night. One developer's public estimate, $1 to $3 a day on DeepSeek or MiniMax against $50 to $130 a day on heavy Opus use, lands on the same shape from the other direction.

Hermes ships the levers for this, and most are off by default. agent.max_turns, the number of tool-calling iterations per turn, is unlimited unless you set it. agent.budget_warning_ratio gives a one-time warning at a fraction of a budget you define. Compression kicks in at 50 per cent of the context limit, and /compress does it on demand; /usage shows what a session has consumed. Subagents are the quiet one: they run their own conversations and hand back only a summary, so a long research goal need not drag its whole history through every call. The ideas behind this are the same as in our piece on context rot in AI agents; the cost angle is just the same problem with a price on it.

MiniMax M3, Claude, or a box under the desk

Which route wins depends less on the model's quality than on how many times a day the agent wakes up. How often would yours? An agent that answers a Telegram message a few times a day is a different economic object from one left on a /goal loop overnight, and the same model can be the cheap choice for one and the expensive choice for the other.

RouteUpfrontLight use (≈50 calls a day)Heavy use (≈500 calls a day)What you give up
MiniMax M3, own key$0Well under $1 a day; cents with cachingAbout $3 a day uncached; under $1 cachedA provider outside the big three; verify data terms
Claude Sonnet 5, own key$0$1 to $2 a day; under $0.20 cachedAbout $19 a day uncached; $2 cachedPrice; the strongest reasoning per call
Nous Portal Plus$20 a monthCovered by the $22 credit at cheap modelsCredits run out; top up or step to SuperYou pay 10% less than list but on a card, not a tap
Local (Ollama, 16–24 GB GPU)The GPUElectricity onlyElectricity only; slower per callModel quality; the hardware bill; your evenings
Managed host$5.99 to $11.99 a monthCovered by included creditsCredit ceiling; check what "included" meansControl over which model and how much

The daily figures are the first table's arithmetic scaled by call count and rounded; a "call" here is one request carrying the roughly 19,000-token average from the measurement. They are not what you will be billed. They are what the price lists and the one available measurement imply, which is the most anyone can honestly say without running your workload.

MiniMax M3 is the interesting entry because its cached-input price, six cents per million, is what makes it cheap for agents specifically. Its uncached input price is the same as Gemini's small model and only a little above GPT-5.6 Luna; the difference shows up over a night of repeated preambles. The Hermes and MiniMax M3 guide covers the model itself and the one-command connection; one detail from the providers documentation is that Hermes's default MiniMax model is still listed as M2.7, which the platform prices at the same $0.30 and $1.20, so switching to M3 is a config line rather than a cost decision. The 1-million-token context is the other reason people pair them, and it is also the reason to watch the 512,000-token tier boundary.

Claude is the opposite bet. Per call it is the most expensive line in the table, and per call it is also the one most likely to finish a multi-step task in fewer calls, which the table cannot show. A session that needs 40 Opus calls to do what takes 120 M3 calls is not sixteen times dearer; it is closer to six, and for some work that is the right trade. The honest position is that nobody outside your own /usage output knows the ratio for your tasks, which is a good argument for starting on the cheap model and keeping the switch command handy. The Claude and DeepSeek profiles have the model-side detail.

Local is where the cost question turns into an ownership question. You run a 26-billion-parameter model on your own card, the tokens are free, the memory and skills files Hermes accumulates live in ~/.hermes on a disk you can hold, and nobody's price list can change under you. What you pay is the hardware, a slower loop, and a model that is a generation behind the hosted ones on hard reasoning. There is a tension here that does not resolve into advice: the free route is the one that costs the most up front, and the metered route is the one where someone else owns the thing that makes the agent clever. Which of those bothers you more is a question about you, not about Hermes. Our AI cost calculator lets you put your own call counts against the current prices, and the model deployment section of the directory lists the serving tools for the local route.

Nous Research, for its part, captures value at the Portal and, since the summer, from investors: the company was reported in July to be raising at a $1.5 billion valuation on the back of an agent it gives away. The software is free, the learning loop is yours, and the thinking is metered by whoever owns the model. That is the arrangement, and once you can see it, the monthly number stops being a surprise.

Everything above is arithmetic on published price lists and one community measurement from an earlier release; the current builds send a different preamble, providers change prices without notice, and the only figure that is true for you is the one in your own /usage report.

Frequently asked questions

Is Hermes Agent free?

The software is, under the MIT licence, and Nous Research's FAQ says you pay only for the model you connect. The tokens are the cost. Connect a local model through Ollama and even those are free, at the price of a graphics card with 16 to 24 GB of memory and the electricity to run it. The install guide walks through the first-run model choice.

How much does Hermes Agent cost per month?

On the one measured workload, an evening of about 200 calls priced at today's lists, a month of daily use runs from roughly $5 on DeepSeek's small model and $7 to $35 on MiniMax M3, depending on caching, up to $59 to $585 on Claude Opus 5. Light, message-a-few-times-a-day use is cents on the cheap models. Nous Portal starts at $20 a month for $22 of credit, and Hostinger's managed box is $5.99 to $11.99 a month with credits included. The AI cost calculator takes your own call count.

Can I run Hermes Agent for free?

Three ways, each with a catch. A local model through Ollama is free per token but needs the hardware. Nous Portal's free tier gives access to free models only. OpenRouter lists a handful of models at $0 per token, with daily request limits. Our advanced LocalAI guide covers the hardware route in depth, and the security guide is worth reading before any free model gets tool access to your machine.

Is MiniMax M3 cheaper than Claude for Hermes?

Per token, by a wide margin: $0.30 in and $1.20 out on the MiniMax platform against $2 and $10 for Claude Sonnet 5 and $5 and $25 for Opus 5, and M3's cached-input rate of $0.06 is what makes the repeated agent preamble cheap. Per finished task the gap narrows if the stronger model needs fewer calls, which only your own usage report can show. The MiniMax M3 guide has the model detail and the connection steps.

Does Hermes work with local models?

Yes. The providers documentation lists Ollama, vLLM, llama.cpp, LM Studio and SGLang, with Ollama needing a context window of at least 64,000 tokens. Ollama's own Hermes page suggests Gemma 4 at around 16 GB of video memory or Qwen 3.6 at around 24 GB. Serving tools for that route are listed under model deployment in the directory.