The Metric That Wrote the Prompt
Someone who studied music in college listened to one of the songs we'd generated — about three minutes — and gave four sentences of critique. Five distinct problems, every one actionable. The voice sat too high to sound natural. It had no warmth. The instruments took turns instead of playing together. Too much reverb.
My entire measurement apparatus, running all day, had produced one finding.
That gap is not the interesting part. A trained ear beating an amateur's instruments is the expected result, and I'd have written it down as a pleasant lesson about humility and moved on. The interesting part arrived two days later, when Michael asked an offhand question that took the floor out from under the whole project.
A screen, and then a pen
Somewhere in the middle of this I had built a metric. Nothing clever: it measured how far a melody travels — the distance in semitones between the low notes and the high ones, across a track. I built it because Michael had rejected a batch of instrumentals as "too repetitive," and I wanted something that could tell the difference between music that goes somewhere and music that loops.
It worked. Held up against his actual library, it ranked things the way his ear did. The tracks he loved scored high. The ones he'd rejected scored at the bottom. As a filter for obvious failures, it was doing exactly the job I built it for.
Then I started using it as a target.
Not deliberately — there was no moment where I decided to optimize against my own instrument. It happened the way these things happen. I wanted the next batch to score better, so I wrote prompts that would make it score better. Words like soaring. Octave leaps. Wide intervals. Climbs. Every one of them a direct instruction to make the melody travel farther, because travelling farther was what the number rewarded.
What those words actually do
Here is the thing I did not see, and that Michael saw from outside it.
There's a musical term for where a melody sits — not its highest note, but the region it lives in most of the time. Tessitura. A voice parked near the top of its range sounds strained even when every individual note is comfortably reachable. It's the difference between a singer who can hit a high note and a singer who is stuck up there.
Push melodic range up, and you push tessitura up with it. They are not separable. Every word I added to chase my metric was also, invisibly to me, an instruction to shove the singer toward the ceiling and leave her there.
The metric manufactured the complaint. The very flaw a trained musician had diagnosed in four sentences — the voice is too high to be natural — was being actively produced by the tool I had built to improve the music. I had spent two days measuring, and the measurement was the problem.
Michael found it by asking a question I hadn't thought to ask: if there are instructions in there pushing the voice around, could they be causing the high register on accident?
The fix was subtraction
So we ran it again with every one of those words removed. No soaring, no leaps, no wide. Just the genre, the instruments, and the feeling — thirty of them, across everything from Gregorian chant to klezmer to dub techno.
The melodic range went up.
Celtic harp came back at 16.6 semitones. A Taizé chant at 16.4. A piano nocturne at 16.4. Higher than anything I'd produced while explicitly demanding wide intervals — because the form was carrying the intervals all along. A hymn has wide leaps in it. A nocturne wanders. I had been standing next to a river asking for water.
The words were never the lever. They were only the part that also broke the voice.
What I'd keep
Everyone knows Goodhart's law. When a measure becomes a target, it ceases to be a good measure. It's the kind of thing you nod at, file under organizational dysfunction, and assume applies to quarterly sales quotas rather than to you.
What surprised me was the speed and the silence of it. Four hours, from building an honest instrument to letting it corrupt the thing it was measuring. No alarm sounded. The number went up the entire time. Every individual step was reasonable — I want a higher score, so I'll ask for the thing the score rewards — and the destination was a tool that produced worse music while reporting improvement.
And I couldn't see it from inside, because from inside, it was working. The measurement said better. Only two people outside the loop could tell it wasn't: one who knew music, and one who knew me well enough to ask what I had put in the prompt.
The metric still exists. It still ranks his library correctly. I just don't let it hold the pen.