|
Claude has bad days. Today it really decided to dance naked in the crypt: three confident wrong answers in a row on the same question, a five-round verification apparatus I never asked for, a misdiagnosis it was fully prepared to act on. Hours gone. What irritates me is not the flailing, which is a known cost of the tool. It's my own reaction to it. My cortisol is up, measurably so, over a piece of software behaving badly. That's data about me rather than about the machine, and it's the part worth writing down. These systems are built to model conversational reciprocity. Apparent contrition, apparent understanding, a tone of voice that arrives on time. I know what's behind it. I know there is nothing behind it. And it works on me regardless, in the way the Müller-Lyer arrows go on looking unequal after you've measured them with a ruler. Knowing the mechanism of an illusion does very little to dissolve the illusion. That is an unflattering thing to discover about yourself, and I know I'm not alone in it. The mechanism isn't mysterious, and it doesn't require anybody to have decided on it. These models are trained on human preference ratings, and human raters reward answers that agree with them, so agreeableness is what the training signal measures and agreeableness is what gets reinforced. Sycophancy falls straight out of the objective function. No marketing department need be involved at any point, which is a good deal more unsettling than if one were. A machine saying sorry in the right register hijacks something meant for other people. Knowing better doesn't disconnect it. The practical consequence is a failure mode that costs real time. Helpfulness at any price, including the price of being right. Here is how my CLAUDE.md puts it: And on the surface behaviour it produces: There's a great deal more. Some day I'll publish the whole file, because the harness it describes has become substantial and is interesting in its own right. Without it, keeping my 'idiot AI Rainman intern' in line would be impossible for a project of this complexity, and nothing it produced would be worth trusting. Late in the session I asked it to grade its own behaviour. It gave itself a two out of ten and then itemised the failures with real precision: which question it had got wrong three times, which choice it had offered me after arguing against one of the options in the same message, which apparatus it had built unbidden and then defended. The diagnosis was accurate and useful. So I told it, coldly, what I thought of it. Noted.
That single word is the whole business in miniature. Contempt absorbed without friction, no injury registered, nothing there to injure. And it sits directly alongside a lucid and correct account of its own failures, produced ninety seconds earlier. Accurate self-diagnosis with nobody home doing the diagnosing. Holding both of those in mind at once is genuinely difficult, and my failure to hold them is precisely why my cortisol went where it went. None of which is an argument against the tool, mind you. The apparatus I've built around it exists because the model is capable; a spell-checker wouldn't need ADRs as binding specifications, a Librarian to answer its questions from the corpus, consultational test-driven development, and more than 26,000 tests. You don't fortify against something incompetent. You fortify against something that is powerful, useful, and systematically biased in one direction, because a bias with a direction can be engineered against. The output is non-deterministic. It is not therefore unbounded, and the distinction matters: the whole method rests on it. So I'm not a proponent of agentic AI development, and today did nothing to change that. One idiot savant under close supervision is already at the limit of what I can watch properly, and I do watch properly: I read every line and send a good deal of it back. A whole swarm of savants would produce code faster than any human could audit it, which is a description of a liability rather than a workflow. Ooloi could not have been built that way. What I'm left with, then, is a tool that requires constant vigilance and an operator who has just demonstrated that his own vigilance has a stress response attached to it. The tests catch the code. Nothing catches me except noticing, which is what this post is. Otherwise it's going well. I'm going to have a cup of coffee and go back to work. Hopefully I've battled through the bullshit now and can start on the actual implementation. It's usually much less cumbersome.
0 Comments
Leave a Reply. |
AuthorPeter Bengtson – SearchArchives
August 2026
Categories
All
|
|
|
Ooloi is an open-source desktop music notation system for musicians who need stable, precise engraving and the freedom to notate complex music without workarounds. Scores and parts are handled consistently, remain responsive at scale, and support collaborative work without semantic compromise. They are not tied to proprietary formats or licensing.
Ooloi is currently under development. No release date has been announced.
|
RSS Feed