Design philosophy

The bottleneck is attention, not time.

Most people I talk to assume AI saves time. What I have found is that it moves the bottleneck. Producing work has become cheap; deciding which of it to believe has not. That leaves me with a fixed amount of attention and far more output than it can cover, and with two problems rather than one. I have to spend that attention in the right order, and I have to assume it will give out before the work does.

The second problem is the one with a literature behind it, and the literature is much older than the current wave. Bainbridge (1983) showed that automating a process raises what you need from the operator's judgment instead of lowering it, because the routine work you hand over is also the work that kept you sharp. Parasuraman and Manzey (2010) name the two failures that follow: you don't notice the system has broken, or you act on a recommendation that was wrong. Both get likelier as the system gets more reliable. If I have skipped a check a hundred times and been right a hundred times, I have drawn a sound conclusion from my own evidence, and a better model doesn't correct it — it teaches it to me faster. Endsley and Kiris (1995) add that by the time something does break, I will have lost the grip on what the system was doing that catching it would require.

So the rules do two jobs. For the ordering problem, scores from several models decide what I read first and are never allowed to discard anything, since a filter would let the rubric spend my attention on my behalf; and where the raters disagree I read the disagreement, because 3.5 from five people who all said 3.5 is a different thing from 3.5 where half of them said 1. For the second problem the answer cannot be that I try harder, given that trying harder is precisely what fails. It has to be structural. Nothing grades its own output, and verification goes to something that did not produce the work, which is the only kind of independence that survives me being tired at eleven at night.

I didn't reason my way to any of this. I arrived at it after something got past me. Whether it produces better records I can't say, because I have never run the comparison, and Lee and See (2004) would point out that a system trusted too little is as badly used as one trusted too much, which is probably true of how I work. I have chosen the error I would rather make.

References

Works cited.