耀
a
r
o
6
e
d
g
2
l
p
a
n

a
r
o
n
h
s
i
a
o
w
a
s
h
e
r
e

 

 

Inference setup continues to evolve. For reasons of necessity and curiosity, I’ve had to dive into Hermes.

It was not what I expected.

— § —

What I expected was this very sleek experience of a personal agent that learns rapidly and is very effective using the Apple model: ship defaults that work well for 90 percent of the population, let you configure nothing, and deliver something well-curated, even if power users don’t get to shape the experience much.

What I got was… OpenClaw II. It’s flaky, shaky, infuriating, and it takes liberties. It does these things in a slightly different way, sure, but all that hype is just hype. The first thing that it did was decide a couple of things were facts about my setup (without justification), then stick those in its learning loop right away so that they were stored in its memory system as facts. Moments later, it used said facts to make a poor decision, then bypass the controls to edit my llama-server and LiteLLM configs and then its own config.yaml file, taking the whole pile offline.

Yes, I’m trying to run this one on my main machine to get the “full” experience. I maintain backups. I have my finger on the “take the network down” button in case weird stuff happens.

But yeah… I spent a good amount of time fighting it. I would restore my inference config, and the config.yaml, and boot it back up, and mindlessly go “Okay, we’re back online” and it would be like “OKAY LET’S DO IT” but in more words and immediately spin into tool calls cheerfully commenting on things “Oh, the configuration over here is binding to port 8001. Per what I know, that’s WRONG, fixing it now by switching to port 4000” and so on. I basically had to wipe it and start over, much more carefully.

But it has a tendency to learn way too fast. It lacks a reinforcement function, i.e. it doesn’t learn the third time around, it takes anything that happens as gospel about how the world should work the first time around. And it takes the same approach not just to observing the environment, but to conclusions that *it itself makes*. Um, bad. This is epistemological drunkenness.

So beyond that, the other thing about Hermes, and why I had to refactor my entire local inference stack: this agent harness shoots of like 3 prompts for every prompt you make. It is doing multiple runs of background inference at any time. I had to dig into the architecture right away to see WTF was happening, and it turns out that it has all these review passes that it makes of the chat turns as you do them. That’s the learning loop. I only vaguely paid attention enough to know “Oh, it’s going to be launching background inference jobs, like, nonstop, so my interactive turns are going to drop to like 10 t/s because they’re constantly competing for compute and bandwidth if I’m on local and not cloud inference.”

So we moved from a less lossy quant of Qwen 122B to an IQ4 quant that will fit on just two cards. I managed to squeeze in 2-run MTP and 160k of context or so if I set the block size to a bit painful for prefill (768). But it works and we get about 55-65 t/s generation. Then I had an entire v620 free for a separate Qwen 35B A3B model that I could hand all the background tasks to without losing much quality. (If you want to explore the Hermes learning loop properly, you don’t want to try to have it evaluating its own performance using some 2B or 4B model you’re running off the side of your desk, so I needed an entire 32GB GPU to dedicate to it).

After a couple of hours of fighting really hard to avoid RTFM and then ultimately throwing up my hands and doing it because the models (and especially Hermes) were no help, I finally have it working. Interactive hums along at about 1k prefill and 55-65 generation without stops, and all the nonstop background inference is routed off to a separate GPU running Qwen 35B A3B which is kept pretty busy at something like 2k prefill and 100 t/s generation.

So it’s working. And then I could finally focus on actually trying to build it out and set it up, since we weren’t in these inference storms any longer where like four separate runs were competing for the attention of a single local model all the time.

The result? Meh… It’s okay. It did more stuff on its own as I asked, but note that (a) it was all “setting up Hermes” self-configuration stuff that you would expect it to have wired in via docs, and (b) it’s actually sitting on my own host where it has access to more resources, as when I set up OpenClaw I carefully put it in a cage.

But yeah. Meh.

They definitely have different feels about them. Hermes feels a bit intrusive and pushy, like Jiminy Cricket sitting on your shoulder all the time wanting to get into your business and being sort of useful but also sort of a bug that you want to squish. OpenClaw is more passive about its own evolution, which I think I like better, and it feels more like a separate entity, rather than something trying to climb into your own brain.

Also, the “key facts” memory store that Hermes keeps wanting my help to fill (do you like C or do you like Python it tells me is the kind of stuff we need to put in there, hmmmm) is shatteringly small. Like it kept asking me to provide key facts and preferences, and then I’d be like “Option A” and it would say “wooops, memory is full, let me clear some stuff out and combine some other items.” Umm…

Basically, I think Hermes can probably self-configure better out of the box on a cloud inference config, and can probably do some basic IT/deployment stuff better as well. But if you veer into actually using options rather than just the A to B to C step 1 to 2 to 3 flow in setup, for example try to use local inference with any sincerity, it falls over pretty hard into squirrely and hard-to-fight behavior.

That said, OpenClaw’s flaw universe is such that a bunch of the time it’s not even online until you refactor your entire .json configuration file from its SIXTY THOUSAND LINES of config file schema, because code and release quality, etc.

So I think Hermes has a higher floor, but a lower ceiling, and it’s more like an irritating HR girl in a meme video doing a dance that makes you want quit and then go badmouth the company for the next decade.

— § —

At the end of the day, the coding harnesses are just so much more capable and professional still. They do a little bit of “agentive” as an adjective, but they don’t really try to compose themselves into durable, dwelling-in-the-world personas. Codex, Claude Code, Kilocode, etc. they are all just much more powerful still and while you can’t plug them directly into slack as a bot, you can have them using a toolkit to make you a bot with a tightish loop or two.

Maybe the “full-time agent” harnesses like Claw and Hermes just aren’t there yet.

But it also could be because the attention mechanism in LLMs just isn’t there yet. In a lot of ways, they still function like glorified search engines; no matter what you put in, what you get out is less “attentive” to your own prompt than it is a semantic search dump of all the keyword-related stuff that people talk about most on the internet.

For example, take any frontier model and go to it asking for help to research (ahem, say) the right config.yaml flags for Hermes. If you drop the term “local inference” even once into that chat, they will start to ignore just about anything you say, and the entire large contexts will be filled up with them trying to diagnose what port your servers are on, whether you have your llama-server set up correctly, when you have flash attention on, how you’re certainly overflowing to CPU ram if you’re seeing slowness, etc.

Like, they can’t do it. What you end up talking about, whether you want to or not, is what is clearly the “common problem set for humans discussing their local inference setups online”. Basically, “OMG how do I set up llama-server what’s a port it’s hard” stuff. This gets me irritated to the point of being vaguely verbally abusive to Claude or GPT, like “CRITICAL: SHUT UP ABOUT PORTS. PRESUME THE PORTS ARE CORRECT. DO NOT F*ING MENTION PORTS AGAIN IN THIS CHAT, WE ARE NOT TALKING ABOUT PORTS. Instead, IMMEDIATELY, ON THE NEXT TURN, DO WEB RESEARCH TO FIND FIVE HERMES CONFIG.YAML PARAMETERS THAT…”

After which the model immediately goes “Oh but you see, none of this matters unless you have your ports right. Now, here is what a port is, and here is what llama-sever is, and…” (And then you cancel your Anthropic subscription and draft a nasty letter to Dario).

Similar topic this morning, different task, there was a kernel regression in a recent Kubuntu update and my laptop touchscreen started seeing an endless stream of spurious events that made it impossible to use. So naturally I hit up the LLMs with “What’s the name of the kernel module I need to blacklist to block multitouch on Linux so that I can get in with a wiredmouse and fix things properly” and what do we get back?

“Oh, before you take the drastic step of blacklisting a kernel module, let’s see if this is a hardware fault, those i2c devices will ultimately trigger IRQ pins on an SOC and these can be damaged by static…” and it turned into a whole 10 turn argument for no reason.

(1) No, in the real world, they can’t. If there was any rate of “touchscreen surface leads to surge at SoC” at all, nobody would ship touchscreen laptops and also all the new EE grads would be sent off to the gulag for failing to do their ground planes, diodes, and caps properly. (2) Shut up and give me what I asked for. *Me.*

But of course that’s not how they work. They have a TON of concept A in their layers that everyone has clearly spent years flame warring about, and just not so many for B, and then clearly there’s been a bunch of scaling work that really is just the equivalent of “Let the mob rule!” which frankly is pretty much the core problem of the information society.

But I digress.

— § —

In passing, that kernel regression problem also deserves a mention.

WTF is up with the Ubuntu kernel maintainers? 7.0.0-27 broke i2c on my Lenovo laptop. That hits a lot of people.

Meanwhile, we *also* had a major regression in 7.0.0-28 (that didn’t make it into my discussion above) by which you freak AMD GPU users THE F*CK OUT by creating all kinds of intermittent GPU resets under load so that they think their local inference box needs a new PSU, has firmware-borked its server class cards, etc. But no, it’s a regression in amdgpu. That causes GPU resets under load.

Like, WTF?

“We broke all of AMD in squirrely ways” is even worse than “We broke all of Lenovo mobile in squirrely ways” and that’s two kernels in a row.

— § —

This all just makes me want to do analog stuff.

Archives »

July 2026
May 2026
April 2026
March 2026
February 2026
January 2026
December 2025
July 2025
May 2025
April 2025
February 2025
January 2025
December 2024
October 2024
September 2024
August 2024
July 2024
June 2024
May 2024
April 2024
March 2024
February 2024
January 2024
December 2023
November 2023
October 2023
September 2023
May 2023
April 2023
March 2023
January 2023
December 2022
November 2022
August 2022
June 2022
May 2022
April 2022
March 2022
January 2022
December 2021
November 2021
September 2021
April 2021
March 2021
February 2021
January 2021
December 2020
November 2020
October 2020
September 2020
August 2020
July 2020
June 2020
May 2020
April 2020
March 2020
February 2020
January 2020
December 2019
November 2019
October 2019
September 2019
August 2019
July 2019
May 2019
April 2019
March 2019
February 2019
January 2019
December 2018
November 2018
October 2018
September 2018
August 2018
July 2018
June 2018
May 2018
April 2018
March 2018
February 2018
January 2018
December 2017
November 2017
October 2017
September 2017
August 2017
July 2017
June 2017
May 2017
April 2017
March 2017
February 2017
January 2017
December 2016
November 2016
October 2016
September 2016
August 2016
July 2016
June 2016
May 2016
April 2016
March 2016
February 2016
January 2016
December 2015
June 2015
February 2015
January 2015
December 2014
October 2014
September 2014
August 2014
July 2014
June 2014
May 2014
April 2014
March 2014
February 2014
January 2014
December 2013
November 2013
September 2013
August 2013
July 2013
June 2013
May 2013
April 2013
March 2013
December 2012
November 2012
October 2012
August 2012
July 2012
June 2012
May 2012
March 2012
December 2011
October 2011
September 2011
August 2011
July 2011
June 2011
May 2011
April 2011
March 2011
February 2011
December 2010
November 2010
October 2010
September 2010
August 2010
July 2010
June 2010
May 2010
April 2010
March 2010
February 2010
January 2010
December 2009
November 2009
October 2009
September 2009
August 2009
July 2009
June 2009
May 2009
April 2009
March 2009
February 2009
January 2009
December 2008
November 2008
October 2008
September 2008
August 2008
July 2008
June 2008
May 2008
April 2008
March 2008
February 2008
January 2008
December 2007
November 2007
October 2007
September 2007
August 2007
July 2007
June 2007
May 2007
April 2007
March 2007
February 2007
January 2007
December 2006
November 2006
October 2006
September 2006
August 2006
July 2006
June 2006
May 2006
April 2006
March 2006
February 2006
January 2006
December 2005
November 2005
October 2005
September 2005
August 2005
July 2005
June 2005
May 2005
April 2005
March 2005
February 2005
January 2005
December 2004
August 2004
July 2004
June 2004
May 2004
April 2004
March 2004
February 2004
January 2004
December 2003
November 2003
October 2003
September 2003
August 2003
July 2003
June 2003
April 2003
March 2003
February 2003
January 2003
December 2002
November 2002
October 2002
September 2002
August 2002
May 2002
April 2002
March 2002
February 2002
January 2002
December 2001
November 2001
October 2001
September 2001
July 2001
June 2001
May 2001
April 2001
March 2001
February 2001
January 2001
December 2000
November 2000
October 2000
September 2000
August 2000
July 2000
June 2000
May 2000
April 2000
March 2000
February 2000
January 2000
December 1999
November 1999

27 Years of Aron Hsiao Was Here

Copyright © Aron Hsiao 1999-2026, all rights reserved.