<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom"><title>Deciphering Glyph</title><link href="https://blog.glyph.im/" rel="alternate"></link><link href="https://blog.glyph.im/feeds/all.atom.xml" rel="self"></link><id>https://blog.glyph.im/</id><updated>2026-09-27T21:27:00-07:00</updated><entry><title>What Would A Serious AI Product Look Like?</title><link href="https://blog.glyph.im/2026/09/serious-ai-product.html" rel="alternate"></link><published>2026-09-27T21:27:00-07:00</published><updated>2026-09-27T21:27:00-07:00</updated><author><name>Glyph</name></author><id>tag:blog.glyph.im,2026-09-27:/2026/09/serious-ai-product.html</id><summary type="html">&lt;p&gt;Every “AI” tool is missing critical features that you would need if
you wanted to do real work with them.  What would those look like, if they
existed?&lt;/p&gt;</summary><content type="html">&lt;body&gt;&lt;p&gt;One of the issues that I have with the current generation of “AI” products is
that they do not appear to take their own premises seriously.  I look at a
plethora of obsequious chatbots claiming to be serious tools for problem
solving, and I think, this is not what a problem-solving tool would look like.&lt;/p&gt;
&lt;p&gt;Even before we get to the tremendous ethical problems with the frontier labs,
it is this impression of their composition &lt;em&gt;as a product&lt;/em&gt; that makes me feel,
constantly, whenever I am interacting with them, that they are less a software
product than that they are a grift, a scam designed to make me feel like I am
interacting with a product that has capabilities that it simply does not, to
try to lull me into a false sense of security that I can trust it.&lt;/p&gt;
&lt;p&gt;The frontier labs are of course the worst offenders, but every criticism here
applies just as much to Ollama, which (if anything, due to the obviously poorer
quality of the available models themselves) needs these features &lt;em&gt;even more&lt;/em&gt;
than the frontier labs do.&lt;/p&gt;
&lt;p&gt;Here, I will set down a few features that might convince me that an LLM-based
product, particularly one focused on research or software development, was
actually serious about helping me do useful things with it.&lt;/p&gt;
&lt;h2 id=make-checking-for-mistakes-a-first-class-feature&gt;Make “Checking For Mistakes” A First-Class Feature&lt;/h2&gt;
&lt;p&gt;This is the biggest issue, and the major reason that I was inspired to write
this post.&lt;/p&gt;
&lt;p&gt;It is a truth universally acknowledged, that AIs cannot reliably provide
information.&lt;/p&gt;
&lt;p&gt;I could cite a ton of news articles and studies about this fact, but there is
no need.  Every single chatbot admits this, up front, in a fine-print
disclaimer as a core part of their user interface.  Gemini says “AI can make
mistakes, so double-check responses”, Claude says “Claude is AI and can make
mistakes. Please double-check responses.&lt;sup id=fnref:1:serious-ai-product-2026-9&gt;&lt;a class=footnote-ref href=#fn:1:serious-ai-product-2026-9 id=fnref:1&gt;1&lt;/a&gt;&lt;/sup&gt;” ChatGPT says “ChatGPT can make
mistakes. Check important info.”.&lt;/p&gt;
&lt;p&gt;Every time I see that last one, I wonder how I’m supposed to know what “info”
is supposed to be “important”.&lt;/p&gt;
&lt;p&gt;All of these warnings are all small, gray text, painfully obviously included as
legalese to push responsibility back onto the user rather than to help with
anything.  This is a &lt;em&gt;core limitation&lt;/em&gt; of all these products.  Checking their
output is a part of the workflow for using them that:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;you &lt;em&gt;absolutely cannot skip or skimp on&lt;/em&gt; without creating risks to yourself
   and whoever you are conveying its output to, and,&lt;/li&gt;
&lt;li&gt;it is &lt;em&gt;very easy to skip or skimp on&lt;/em&gt; and you are encouraged at every turn
   to do so, because “just trust the output” is one of the quickest ways to
   save time.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;A chatbot product that took this weakness seriously, as an actual consideration
for using it, would put a checkbox next to &lt;em&gt;every claim&lt;/em&gt; in its output.  It
would be a 2-column worksheet, where you’ve got the LLM output in the first
column, and next to it, human notes in the second column, explaining what work
went into checking this claim, and a &lt;em&gt;big&lt;/em&gt; checkbox that you would only check
off after you believe you’d checked its claims thoroughly enough.&lt;/p&gt;
&lt;p&gt;Coding assistants would need to have some version of this as well.  Right now,
this is pushed off into code review, which means it is a dark pattern which
subtly encourages the “author”&lt;sup id=fnref:2:serious-ai-product-2026-9&gt;&lt;a class=footnote-ref href=#fn:2:serious-ai-product-2026-9 id=fnref:2&gt;2&lt;/a&gt;&lt;/sup&gt; to offload this work to their code reviewer
without ever looking.  Once again, “it’s probably fine, I don’t need to check”
is the quickest way to save time and churn out those PRs faster.&lt;/p&gt;
&lt;p&gt;It might even be useful for coding harnesses to have some affordance for
checking code before it even runs tests.  &lt;a href="https://claude.com/blog/agentic-coding-is-straining-ci-heres-how-we-scaled-test-impact-analysis-at-anthropic"&gt;As the vendors themselves have
admitted&lt;/a&gt;,
it’s not just expensive to burn tokens on your “AI”, you also end up burning
far more &lt;em&gt;compute&lt;/em&gt; on the AI.  Being able to check your diffs before sending
them over to uselessly exhaust your testing compute cluster would be useful.&lt;/p&gt;
&lt;p&gt;If your product tells me that it makes mistakes and I must be the one to check
for the mistakes, but then gives me &lt;em&gt;zero tools&lt;/em&gt; to check for mistakes, I
cannot take it seriously.&lt;/p&gt;
&lt;h3 id=more-citations-to-check-and-more-details&gt;More Citations to Check, And More Details&lt;/h3&gt;
&lt;p&gt;Most chatbots prefer to give an answer, rather than a citation.  In my own
personal use, I find that when asked to provide a list of citations with
clearly marked sources for each one, they will appear to “get bored” halfway
through the list and simply stop including citations at some point.&lt;/p&gt;
&lt;p&gt;When the bots include citations at all, present them as inline annotations that
say nothing but the domain name of the search result, in a font so small that
it’s barely legible, and an equally indecipherable icon that is fewer than 16
pixels on a side.&lt;/p&gt;
&lt;p&gt;This is backwards.&lt;/p&gt;
&lt;p&gt;Now, I am aware that these citations do come from somewhere, and in an attempt
to reduce hallucinations, all of the major providers support some form of
“grounding”&lt;sup id=fnref:3:serious-ai-product-2026-9&gt;&lt;a class=footnote-ref href=#fn:3:serious-ai-product-2026-9 id=fnref:3&gt;3&lt;/a&gt;&lt;/sup&gt;, and that those little barely-readable citation links are
referencing actual structures in the
&lt;a href="https://en.wikipedia.org/wiki/Retrieval-augmented_generation"&gt;RAG&lt;/a&gt; pipeline
and not just potentially-hallucinated tokens, but I’m not talking about the
underlying machinery in the model, I’m talking about the &lt;em&gt;presentation to the
user&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;Plus, regardless of whether a snippet of text came from a RAG query, we know
that LLMs can &lt;a href="https://www.computerworld.com/article/4059383/openai-admits-ai-hallucinations-are-mathematically-inevitable-not-just-engineering-flaws.html"&gt;never provide an authoritative
result&lt;/a&gt;;
it’s a fundamental limitation of the technology.  They can still garble the
results of RAG as much as they can misrepresent any other training data.  This
means that it must never present its results &lt;em&gt;as&lt;/em&gt; authoritative.&lt;/p&gt;
&lt;p&gt;If you ask an AI to do research queries, every result should be presented as a
&lt;em&gt;list of citations&lt;/em&gt;.  Moreover, the presentation should display each citation as a
large object of in its own right, with clearly identified metadata, including
not just the site where it was found but its publication date and, if possible,
the name of the author.  The literal, unmodified quotation (not from RAG, not a
summary: a quotation extracted with a regular program and not an LLM) should be
front-and-center, larger than any AI-generated text.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;If&lt;/em&gt; the AI product wants to editorialize or summarize (which should not always
be necessary!), the AI-generated text should be presented as small text
underneath the citation that has been found, de-emphasized as much as the
disclaimer is right now, at the very least until the user has verified that the
summary is accurate.  Perhaps, for a research project, a “did you read the
citation” checkbox might even be helpful.&lt;/p&gt;
&lt;p&gt;If your product openly tells me that it will scramble, misrepresent, or omit
its citations in its summaries, and I must read the original human-authored
citations to be sure, but then gives me no tools to track my reading of those
citations or even any way to &lt;em&gt;find&lt;/em&gt; them, I cannot take it seriously.&lt;/p&gt;
&lt;h2 id=no-first-person-output-no-apologies&gt;No First-Person Output, No Apologies&lt;/h2&gt;
&lt;p&gt;There is no reason for a software development or research tool to use
first-person language to describe itself.  They should not do so.  In fact
&lt;a href="https://www.theatlantic.com/ideas/2026/09/meta-settlement-social-media-addiction-youth/688567/"&gt;they should not be &lt;em&gt;allowed&lt;/em&gt; to do
so&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;There is also no reason that they should ever &lt;em&gt;apologize&lt;/em&gt;.  &lt;a href="https://link.springer.com/article/10.1007/s43681-025-00800-x"&gt;It is a waste of
everyone’s time&lt;/a&gt;;
it’s a waste for the chatbot to generate the apology, it’s a waste for the user
to read the apology, and it’s a waste for the user to respond to the apology.
Yet they unfailingly do this upon every correction.&lt;/p&gt;
&lt;p&gt;The vendors of these tools &lt;a href="https://openai.com/index/strengthening-chatgpt-responses-in-sensitive-conversations/"&gt;know that they are routinely causing mental-health
crises&lt;/a&gt;.
In response, they have added non-functional “guard rails” that can still, in
2026, &lt;a href="https://www.techzine.eu/blogs/security/141882/chatgpt-easily-bypasses-its-own-guardrails-all-llms-are-inherently-unsafe/"&gt;easily be
bypassed&lt;/a&gt;.&lt;sup id=fnref:4:serious-ai-product-2026-9&gt;&lt;a class=footnote-ref href=#fn:4:serious-ai-product-2026-9 id=fnref:4&gt;4&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;A product seriously interested in helping with productivity would correct this
&lt;em&gt;glaringly&lt;/em&gt; obvious flaw, focus on the task at hand, and stop emitting useless
verbiage.&lt;/p&gt;
&lt;p&gt;In the previous two sections, I tried to focus on ways in which the harness
would be constructed differently even if the LLM technology is fundamentally
impossible to improve; in this case, I have to assume that the labs have &lt;em&gt;some&lt;/em&gt;
control over the model itself.  But unless they are truly incapable of
influencing their output (and all their “benchmarks” and “capabilities” seem to
indicate that they can control it very tightly) they ought to be building
models that are much less verbose.&lt;/p&gt;
&lt;h2 id=more-non-natural-language-user-interfaces&gt;More Non-Natural-Language User Interfaces&lt;/h2&gt;
&lt;p&gt;Although natural language could &lt;em&gt;hypothetically&lt;/em&gt; be a powerful interface for
interacting with a computer system, the practical upshot of LLM natural
language interfaces is that these interfaces are imprecise and repetitive, full
of &lt;a href="https://arxiv.org/abs/2602.11988"&gt;superstitions masquerading as “best
practices”&lt;/a&gt;.  The inputs are a mess and the
resulting outputs are a mess.&lt;/p&gt;
&lt;p&gt;The general way of addressing this unstructured mess is to allow the chatbot to
&lt;em&gt;directly take action&lt;/em&gt; in response to the user’s input; in other words to
supply it with “tools” via an MCP server.  But again, this is backwards.  If we
cannot even express our intent clearly in the first place, why are we trusting
this system to take potentially destructive and harmful actions on our behalf?&lt;/p&gt;
&lt;p&gt;Instead, I would expect a product that was seriously invested in helping me
accomplish specific tasks, to have user interfaces specific to those tasks.  Is
it supposed to be able to be a security scanner that can discover OWASP top 10
bugs in a codebase? Have a button for that.  Build that functionality into your
harness, train it directly into the model, use smaller models that can satisfy
that functionality more effectively than throwing it at the planet-sized brain
of Fable or whatever.&lt;/p&gt;
&lt;p&gt;I’m aware that there are small software startups that do &lt;em&gt;something&lt;/em&gt; like this,
but they are bolted on to the side of the main model providers’ APIs, not
integrated into the core of the product and not using their own models and AI
systems to achieve consistent and repeatable results.&lt;/p&gt;
&lt;h2 id=strong-data-provenance-indicators&gt;Strong Data Provenance Indicators&lt;/h2&gt;
&lt;p&gt;Chatbots produce data tables pulled from websites, from APIs, from MCP tools or
from summarizing and scrambling the user’s input.  In order to provide the
illusion of a seamless interface, this data is presented in-line regardless of
where it comes from.  But some of these outputs are produced mechanically via
regular old API calls, for example, from the result of calling a tool or
querying a website, but presented uniformly.&lt;/p&gt;
&lt;p&gt;But there is a huge difference between an authoritative data source being
inlined as part of a chatbot conversation, being treated as &lt;em&gt;input&lt;/em&gt; by the
chatbot, and some ad-hoc hallucinated data being treated as &lt;em&gt;output&lt;/em&gt; of the
chatbot.&lt;/p&gt;
&lt;p&gt;If a product is trying to help me make accurate, empirically-grounded,
data-driven decisions, the source of the data is critical.&lt;/p&gt;
&lt;p&gt;Integrated into the “check for mistakes” and “verify citations” workflow I
described above, there’s a necessary “verify data programmatically” pass as
well; to have tools that will treat portions of the output as a &lt;em&gt;regular
spreadsheet&lt;/em&gt;, allowing regular-old computer arithmetic to verify things and
&lt;em&gt;showing where such arithmetic was used, and how&lt;/em&gt;.&lt;/p&gt;
&lt;h2 id=better-user-control-of-reproducibility&gt;Better User Control of Reproducibility&lt;/h2&gt;
&lt;p&gt;Anyone familiar with the technical specifics of LLMs will know that they have a
variable called
“&lt;a href="https://www.ibm.com/think/topics/llm-temperature"&gt;temperature&lt;/a&gt;” which
controls the degree of randomness that the LLM uses to produce its outputs.
But most users don’t know this, because it isn’t exposed as part of the user
interface by default.&lt;/p&gt;
&lt;p&gt;This leads to a subjective impression that you asked ChatGPT, and you got
ChatGPT’s authoritative answer.&lt;/p&gt;
&lt;p&gt;You can’t just set the temperature to zero and still get useful results - I am
aware that it does more than just scramble the output at random, and there are
perhaps good reasons that simply exposing &lt;em&gt;just&lt;/em&gt; a temperature setting would
not be that useful to users.  But if we followed some more of my earlier
recommendations for making more structured UI elements to solve specific
problems rather than having long back-and-forth chats where each refinement
depends on the previous response, perhaps those elements could also re-play the
process so that users can see how reliable the bot is at a particular task and
develop a sense of how the stochastic nature of the process actually affects
it.&lt;/p&gt;
&lt;p&gt;Similarly, if a user is trying to solve the same problem repeatedly with a
chatbot, and the chatbot &lt;em&gt;product&lt;/em&gt; has numerous computational tools that don’t
really have anything to do with the LLM, such as deterministic data-processing
tools, then having a way to freeze the non-deterministic parts of the
transcript but re-populate a particular data frame with updated information and
fork / continue the conversation from there would be a way to avoid introducing
pointless additional randomness when you already know what tool you’re trying
to use.&lt;/p&gt;
&lt;p&gt;The fact that every conversation is presented as this flat chat prompt that
doesn’t let me interact with any of the widgets that were previously produced
except through more chatting, really makes me feel like the whole product is
just doing predatory social-media style “increase time on site” optimization,
just trying to lure me into further repetitive and unreliable chats, rather
than letting me get in, solve my problem, and get out.&lt;/p&gt;
&lt;h2 id=context-visibility&gt;Context Visibility&lt;/h2&gt;
&lt;p&gt;Managing the LLM context is &lt;em&gt;the&lt;/em&gt; ongoing challenge facing organizations that
are trying to use “agentic” workflows.  Filling up the context with too much
information &lt;a href="https://www.trychroma.com/research/context-rot"&gt;causes well-known
problems&lt;/a&gt;. In response,
advanced LLM users attempting to solve larger problems must break up very long
prompts into “skills”, give access to lengthy information via “tools”, and
delegating sub-problems to “sub-agents” rather than simply extending a single
prompt indefinitely.&lt;/p&gt;
&lt;p&gt;All of these strategies have flaws, because even on the largest models,
compared to the breadth and depth of knowledge-work problems, LLM contexts are
quite small.&lt;/p&gt;
&lt;p&gt;And yet, none of these products will &lt;em&gt;show the context to the user&lt;/em&gt; by default.
There are third-party addons that can show you a &lt;a href="https://www.ai-toolbox.co/ai-toolbox-chatgpt-features/chatgpt-context-window-meter-2026"&gt;simple progress
bar&lt;/a&gt;
but for addressing the premier engineering difficulty with this technology,
that is below the bare minimum.&lt;/p&gt;
&lt;p&gt;This lack of visibility means that almost all of the tools for extending the
context are flying blind.  Rather than responding meaningfully to a full
context, everyone just kind of guesses how much state they need by guessing and
trying over and over again with progressively more elaborate skill and sub-agent
layouts.  Even managing context compaction ends up being an &lt;a href="https://platform.claude.com/cookbook/tool-use-automatic-context-compaction"&gt;advanced
API-driven
workflow&lt;/a&gt;&lt;sup id=fnref:5:serious-ai-product-2026-9&gt;&lt;a class=footnote-ref href=#fn:5:serious-ai-product-2026-9 id=fnref:5&gt;5&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;A serious product that was trying to help the user understand would not only
show “available context” but explain the impact of context compactions, make it
easier to see harness-generated prompts, and so on.  This would be a
first-class feature, combined with the aforementioned reproducibility / replay
tools, would allow users to do real experiments to develop an understanding
about how to make good use of the context window.&lt;/p&gt;
&lt;h2 id=a-sandbox-that-actually-works&gt;A Sandbox That Actually Works&lt;/h2&gt;
&lt;p&gt;I’ve been focused on the chatbot interface here because it is the most
&lt;em&gt;immediately&lt;/em&gt; egregious upon looking at the UI.  But the “agentic loop” tools
used for coding are equally dangerous, if not more so.  Coding tools
&lt;a href="https://aishippingblog.com/p/how-i-dropped-our-production-database"&gt;keep&lt;/a&gt;
&lt;a href="https://cybersecuritynews.com/claude-code-agent-file-deletion/"&gt;destroying&lt;/a&gt;
&lt;a href="https://x.com/lifeofjer/article/2048103471019434248"&gt;everyone’s&lt;/a&gt;
&lt;a href="https://quasa.io/media/when-cursor-wiped-a-user-s-pc-a-cautionary-tale-of-ai-overreach"&gt;data&lt;/a&gt;,
over the course of &lt;em&gt;years&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;These catastrophic incidents that become front-page news are relatively rare
compared to the amount of coding-agent use out there.  But they also aren’t the
only kind of sandbox violation.  Coding models will so routinely edit test code
instead of the system under test that there are “pro tips” articles all over
the web giving you the &lt;a href="https://claudecodetips.com/en/guide/pitfalls/39"&gt;flawed
advice&lt;/a&gt; to simply &lt;em&gt;ask&lt;/em&gt; the
agent not to cheat.  News write-ups of the catastrophic incidents themselves
will also offer glib and wrong advice, like “use a docker container”.  That
might prevent it from literally deleting your operating system, but it won’t
prevent it from destroying all the local work you have in your codebase (it
needs access to a checkout, after all!)&lt;/p&gt;
&lt;p&gt;There is a flurry of activity in the infosec space where people are rushing to
plug the gaps left by these coding harnesses.  Everyone’s got their own version
of an &lt;a href="https://www.truefoundry.com/blog/mcp-tool-approvals-explained"&gt;MCP&lt;/a&gt;
&lt;a href="https://microsoft.github.io/agent-governance-toolkit/tutorials/07-mcp-security-gateway/"&gt;approval&lt;/a&gt;
&lt;a href="https://github.com/thehairlesshorseman/mcp-approval-gateway"&gt;gateway&lt;/a&gt; where
you can optionally place a proxy between your agent and your production
infrastructure.&lt;/p&gt;
&lt;p&gt;In the best case, though, all these mitigations and proxies and prompts simply
turn the user into an &lt;a href="https://en.wikipedia.org/wiki/Drinking_bird#cite_ref-26"&gt;auto-approval
automaton&lt;/a&gt;, hitting Y,
Y, Y, Y over and over again, until you finally are driven mad and hit “yes to
all”, turn on full-auto mode and submit yourself to the void.  With nothing
between your personal vigilance and disaster, there are no workflows left
beyond decrementing your own vigilance until there’s nothing left and then
hoping the disaster never arrives.&lt;/p&gt;
&lt;p&gt;The fact that &lt;em&gt;some&lt;/em&gt; mitigations exist that can be deployed by extra-cautious
users does not change the fact that “agentic coding” is an unsafe-by-default
technology deployed without concern or guidance.  Every frontier lab has tied a
spring-loaded shotgun to a dog; the fact that dog owners can publish thoughtful
blog posts explaining how you can teach your dog the basics of gun safety or
how you can have your dogs play in a bullet-proof room does not mitigate the
fact that the product should not have been allowed in the first place, nor
should it continue to exist without VERY strong security controls.&lt;/p&gt;
&lt;p&gt;I might believe that a frontier lab were seriously interested in providing
developers with a useful tool if they shipped something that had safety built-in.&lt;/p&gt;
&lt;p&gt;That means tools &lt;em&gt;in the harness&lt;/em&gt;, detached from any LLM, independent of the
prompt, that could:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;sandbox all filesystem operations and strictly limit ANY deletions outside of
  specified scopes, regardless of operating system,&lt;/li&gt;
&lt;li&gt;enforce snapshotting of the entire repo on every operation for easy rollbacks
  and minimal lost work,&lt;/li&gt;
&lt;li&gt;remove &lt;a href="https://cybernews.com/security/claude-code-auto-mode-malware-vulnerability/"&gt;the disaster of “auto
  mode”&lt;/a&gt;
  (not to mention nonsense like &lt;code&gt;--dangerously-skip-permissions&lt;/code&gt;) entirely, and&lt;/li&gt;
&lt;li&gt;carefully consider a structure for presenting plans to the user where, rather
  than provoking immediate alert fatigue by asking for checks on every action,
  make structured plans which can be submitted to the user as a group of
  actions and reviewed and approved as a batch.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In the same way that I suggested above that research-based tasks should have a
way of re-issuing prompts to determine how reproducible a result is, or whether
other sources might be found, agent-based tasks should have a way of being
executed against mock services for popular APIs, so that the verification can
match both on the front-end (review the plan for making the API calls before
they’re executed) and the back end (review the API calls that were issued to
the mock service and verify that they matched).&lt;/p&gt;
&lt;p&gt;Instead, the frontier labs provide us products that are disasters out of the
box, give us “best practices” to build massive and elaborate, as well as
incomplete and error-prone, security perimeters of our own design.  Then they
blame “operator error” when it inevitably goes wrong.  I cannot believe that
these design choices are intended to help us be productive.&lt;/p&gt;
&lt;h2 id=bonus-human-processes&gt;Bonus: Human Processes&lt;/h2&gt;
&lt;p&gt;Organizations &lt;em&gt;deploying&lt;/em&gt; AI also frequently come across as unserious, for
similar reasons.  In 2023, naive exuberance could perhaps be forgiven.  But
today, as we near the close of 2026, there are several well-known problems,
that have been extremely well-covered in the press.  None of these things
should be surprising, but most orgs deploying these tools are still just
letting them rip and hoping it all works out.&lt;/p&gt;
&lt;p&gt;Organizations deploying these tools would need at least three kinds of major
modifications to their internal processes, if they wanted to be serious about
using them safely:&lt;/p&gt;
&lt;h3 id=1-shift-rotations-to-prevent-vigilance-decrement&gt;1. Shift Rotations to Prevent Vigilance Decrement&lt;/h3&gt;
&lt;p&gt;There have been several high-profile incidents where software developers’
gradual acquiescence to accepting LLM output have lead to serious economic
consequences for the companies deploying them, perhaps best typified by
Amazon’s “&lt;a href="https://www.businessinsider.com/amazon-tightens-code-controls-after-outages-including-one-ai-2026-3"&gt;millions of lost
orders&lt;/a&gt;”
due to a gradual decay of their engineering processes from LLM use.&lt;/p&gt;
&lt;p&gt;These outages, and other AI-related failures, are due to the difficulty of
maintaining focus on the same problems.  In other words, as I described above,
&lt;a href="https://www.csoonline.com/article/4225613/aviation-solved-the-vigilance-problem-ai-just-gave-security-a-worse-one.html"&gt;vigilance
decrement&lt;/a&gt;
is a constant problem, because AI outputs are &lt;em&gt;most often&lt;/em&gt; correct, but
continue to be incorrect in surprising and non-intuitive ways.  As I have
&lt;a href="https://blog.glyph.im/2026/03/what-is-code-review-for.html"&gt;previously written&lt;/a&gt;, you cannot trust
yourself to catch every bug with code review, and LLM output.&lt;/p&gt;
&lt;p&gt;Aviation, for example, has &lt;em&gt;very strict rules&lt;/em&gt; around &lt;a href="https://www.ecfr.gov/current/title-14/chapter-I/subchapter-G/part-117"&gt;rest
requirements&lt;/a&gt;.
There is also a specific rule that &lt;a href="https://www.ecfr.gov/current/title-14/chapter-I/subchapter-G/part-135/subpart-B/section-135.99"&gt;“No certificate holder may operate an aircraft
without a second in command if that aircraft has a passenger seating
configuration, excluding any pilot seat, of ten seats or
more.”&lt;/a&gt;.
Other safety-critical professions have similar rules.&lt;/p&gt;
&lt;p&gt;And yet, even in the age of the supposed “AI revolution”, most software teams
are still assigning every engineer a full feature load, not planning for any
rest, and telling people to review code whenever they happen to have some “free
time”.&lt;/p&gt;
&lt;p&gt;Maintenance of vigilance has to be your top priority.  Regular, scheduled,
&lt;em&gt;inviolable&lt;/em&gt; rest periods where people do work without AI assistance, and are
not exposed to any AI output for review or otherwise, would be crucial in order
to stay mentally sharp enough.&lt;/p&gt;
&lt;p&gt;The tools themselves should have this sort of thing built in.  The
mistake-review process described above should have a periodic spot-check mode
where a second reviewer periodically reviews a chatbot log, doing their own
independent verification of claims, to see if they spot the same errors.  This
could provide a feedback loop to determine how much rest is necessary to
maintain continuous attention and actually spot hallucinations.&lt;/p&gt;
&lt;h3 id=2-skill-practice-to-prevent-skill-loss&gt;2. Skill Practice To Prevent Skill Loss&lt;/h3&gt;
&lt;p&gt;It is also well-known that AI use leads to AI reliance, and AI reliance leads
to &lt;a href="https://www.psychologytoday.com/us/blog/the-algorithmic-mind/202603/adults-lose-skills-to-ai-children-never-build-them"&gt;skill
loss&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;I like to use the analogy to dockworkers at a seaport&lt;sup id=fnref:6:serious-ai-product-2026-9&gt;&lt;a class=footnote-ref href=#fn:6:serious-ai-product-2026-9 id=fnref:6&gt;6&lt;/a&gt;&lt;/sup&gt; adopting automation.&lt;/p&gt;
&lt;p&gt;If you employ dockworkers to load and unload ships all day long, they are going
to be getting tons of exercise.  They will be able to lift heavy objects on
demand, whenever.  They might have plenty of health problems and injuries from
this type of work, but “lack of exercise” will not be a problem.&lt;/p&gt;
&lt;p&gt;With the development of standardized container ships and mechanized cranes, you
are going to be changing their job description substantially: now they mostly
spend all day sitting in a small cubicle moving a control lever back and forth,
not lifting heavy stuff.  They will get worse at lifting heavy objects.&lt;/p&gt;
&lt;p&gt;In this analogy however, the cranes are not all that reliable.  We know they
break, and they drop their payloads sometimes, and the stuff needs to be
manually moved.  But this only happens a few times a week, at most.  If you
need whoever is driving the crane to be able to jump out at any moment and
still move stuff around manually, then you need to make an affordance for that.
You need to give them &lt;em&gt;time&lt;/em&gt; to go to the gym and do some lifting for practice,
or every crane failure is going to be a major emergency.&lt;/p&gt;
&lt;p&gt;An organization doing an AI transformation would also need a massive increase
to learning &amp;amp; development budget, both in terms of resources and in terms of
schedule.  If your people are going to lose skills because they’ve lost regular
practice in the incidental course of doing their duties, then they are going to
need &lt;em&gt;deliberate&lt;/em&gt;, intentional, non-incidental practice of those skills to keep
them sharp.&lt;/p&gt;
&lt;p&gt;But rather than trying to accommodate new workflows and give time for people to
adjust, most AI mandates are simply dropped on workers like a ton of bricks,
with no time to adapt and no affordance for maintaining their skills.  Operate
the crane and stay fit and healthy and ready to switch back to manual lifting
at any time and then get back in the crane cockpit right afterwards.  Don’t
mess up.&lt;/p&gt;
&lt;p&gt;Then an accident happens and everyone is surprised, as if this process weren’t
practically &lt;em&gt;designed&lt;/em&gt; to produce a terrible result.&lt;/p&gt;
&lt;h3 id=3-mental-health-resources-to-deal-with-mental-health-risks&gt;3. Mental Health Resources to Deal with Mental Health Risks&lt;/h3&gt;
&lt;p&gt;AI psychosis often begins with &lt;a href="https://pulitzercenter.org/stories/ai-psychosis-mental-health-crisis-21st-century"&gt;practical
problem-solving&lt;/a&gt;,
and beyond that, it can start &lt;a href="https://www.forbes.com/sites/bryanrobinson/2025/11/02/ai-psychosis-at-work-mental-health-experts-express-concerns/"&gt;specifically at
work&lt;/a&gt;.
Not to mention the more pedestrian condition of &lt;a href="https://www.cnn.com/2026/03/13/business/ai-brain-fry-nightcap"&gt;“AI brain
fry”&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;If you are mandating your employees to use a hazardous tool that may seriously
and &lt;em&gt;directly&lt;/em&gt; damage their mental health, you need trainings and resources.
You need in-house therapists and you need to be making sure to check in with
people actively to make sure that this is not happening.&lt;/p&gt;
&lt;p&gt;Again, the tool itself ought to have some way of dealing with this.  An
&lt;a href="https://lifehacker.com/tech/why-chatgpt-is-now-reminding-you-to-take-break"&gt;occasional “take a
break”&lt;/a&gt;
popup is easily dismissed; they need a user-visible AI &lt;a href="https://en.wikipedia.org/wiki/Film_badge_dosimeter"&gt;personal
dosimeter&lt;/a&gt; so you can see
your &lt;em&gt;cumulative&lt;/em&gt; usage over time.&lt;/p&gt;
&lt;p&gt;I don’t even know if “usage over time” is a sufficient metric to gauge risk.
Maybe if your work chatbot start to talk about
&lt;a href="https://medium.com/@ijin2504/from-tool-to-self-the-7-resonance-phases-of-gpt-5fa6954aeb49"&gt;resonance&lt;/a&gt;
too much, unless you literally work as an acoustic engineer, that should be
flagged for someone.&lt;/p&gt;
&lt;p&gt;We are, again, years into dealing with these tools, and we know these risks
exist.  Yet no serious mitigations are provided.  Not even any way of measuring
the risk exposure.&lt;/p&gt;
&lt;h3 id=and-more&gt;And More&lt;/h3&gt;
&lt;p&gt;There are also many other risks associated with the technology.  There are
&lt;a href="https://arstechnica.com/tech-policy/2026/07/openai-faked-inability-to-search-training-data-hid-billions-of-logs-nyt-says/"&gt;intellectual property
risks&lt;/a&gt;
with the foundation models, due to recklessness with their training data.
There are &lt;a href="https://www.cnbc.com/2026/09/24/oracle-data-center-force-majeure.html"&gt;existential financial
risks&lt;/a&gt;
associated with the infrastructure build-out.  The extent to which most “open”
models are &lt;a href="https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-251a"&gt;simply
derivatives&lt;/a&gt;
of frontier models is an open question.&lt;/p&gt;
&lt;h2 id=what-i-think&gt;What I Think&lt;/h2&gt;
&lt;p&gt;If any &lt;em&gt;one&lt;/em&gt; of these things were regularly overlooked by AI vendors or users,
that would be a totally normal product oversight.  Room for improvement for the
next version, but nothing catastrophic.&lt;/p&gt;
&lt;p&gt;Shipping without &lt;em&gt;any&lt;/em&gt; of them doesn’t seem like lean product management, it
seems like a careless attitude towards risk and a product design philosophy
oriented entirely towards short-term demos, with no regard for how to realize
actual productivity gains.&lt;/p&gt;
&lt;p&gt;Furthermore, being available for &lt;em&gt;years&lt;/em&gt; without anything like these features,
despite hundreds of incidents demonstrating the risks, with hundreds of
billions of dollars of funding, makes it seem to me like if they &lt;em&gt;were&lt;/em&gt; to add
all the features that would make their product actually safe and hypothetically
useful, these features would reveal that it is actually not an improvement to
productivity.&lt;/p&gt;
&lt;p&gt;In the &lt;a href="https://blog.glyph.im/2025/08/futzing-fraction.html"&gt;year&lt;/a&gt; since I first wrote about
measuring the cost/benefit ratio of AI, I have heard from numerous people who
have shown this to management to try to illustrate why their AI initiatives —
like &lt;a href="https://ludic.mataroa.blog/blog/ai-mania-is-eviscerating-global-decision-making/"&gt;almost all AI
initiatives&lt;/a&gt;
— were either failing or burning out their engineers.&lt;/p&gt;
&lt;p&gt;I’ve also heard from lots of people that have told me that it’s obviously
useful and they don’t need to measure so carefully, because they are getting
lots of work done that they couldn’t have otherwise.&lt;sup id=fnref:7:serious-ai-product-2026-9&gt;&lt;a class=footnote-ref href=#fn:7:serious-ai-product-2026-9 id=fnref:7&gt;7&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;I have yet to hear from a &lt;em&gt;single&lt;/em&gt; person who has said “yeah, we measured
according to your methodology&lt;sup id=fnref:8:serious-ai-product-2026-9&gt;&lt;a class=footnote-ref href=#fn:8:serious-ai-product-2026-9 id=fnref:8&gt;8&lt;/a&gt;&lt;/sup&gt;, and it turns out that our AI work is going
great and that our ratio is 0.75”.&lt;/p&gt;
&lt;p&gt;Obviously, I cannot say for sure why this is; absence of evidence is not
evidence of absence.  But at this point I think the null hypothesis is that AI
tools provide, in aggregate, zero value.  They make mistakes too often, and the
externalities they produce are so bad and so difficult to control that even
before we get to the places where they are &lt;a href="https://www.scientificamerican.com/article/spacexais-colossus-data-center-is-fueling-a-massive-backlash/"&gt;just physically poisoning
people&lt;/a&gt;,
even the negative effects on their direct users end up cancelling out whatever
benefit to they provide to their organizations.&lt;/p&gt;
&lt;p&gt;If I were wrong, then including tools to &lt;em&gt;measure&lt;/em&gt; an AI’s effectiveness &lt;em&gt;at
the tasks their users are actually trying to accomplish&lt;/em&gt;, rather than
&lt;a href="https://www.makeuseof.com/ai-benchmark-numbers-are-meaningless-heres-what-to-look-for-instead/"&gt;meaningless
benchmarks&lt;/a&gt;,
would show big productivity gains.  The frontier labs would be champing at the
bit to add such features, and crowing about their fantastic results.&lt;/p&gt;
&lt;p&gt;I think the labs know that if they did that, it would present a grim picture to
their users. Such tools would let their users see that it’s making mistakes
much more often than they realized, that they’re spending much more time with
it than they want to be, and that it’s just generally not fit for purpose.&lt;/p&gt;
&lt;p&gt;If they prove me wrong by adding in all of these safety mechanisms, and in the
process, they make all of their AI technology less harmful, I’ll be thrilled to
be debunked.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=acknowledgments&gt;Acknowledgments&lt;/h2&gt;
&lt;p class=update-note&gt;Thank you to &lt;a href="/pages/patrons.html"&gt;my patrons&lt;/a&gt; who are supporting my writing on
this blog.  If you like what you’ve read here and you’d
like to read more of it, or you’d like to support my &lt;a href="https://github.com/glyph/"&gt;various open-source
endeavors&lt;/a&gt;, you can &lt;a href="/pages/patrons.html"&gt;support my work as a
sponsor&lt;/a&gt;!&lt;/p&gt;
&lt;div class=footnote&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=fn:1:serious-ai-product-2026-9&gt;
&lt;p id=fn:1&gt;It is also interesting that for the next section, &lt;em&gt;sometimes&lt;/em&gt; it seems
that Claude’s disclaimer is “Please double-check cited sources.” instead. &lt;a class=footnote-backref href=#fnref:1:serious-ai-product-2026-9 title="Jump back to footnote 1 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:2:serious-ai-product-2026-9&gt;
&lt;p id=fn:2&gt;... by which I mean the “prompter”, since authorship is not what’s
happening here. &lt;a class=footnote-backref href=#fnref:2:serious-ai-product-2026-9 title="Jump back to footnote 2 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:3:serious-ai-product-2026-9&gt;
&lt;p id=fn:3&gt;&lt;a href="https://claude.com/blog/introducing-citations-api"&gt;Claude has the “citations
API”&lt;/a&gt;, &lt;a href="https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/grounding/overview"&gt;Google has
various different kinds of “grounding” against its own
APIs&lt;/a&gt;, and I guess Microsoft can &lt;a href="https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/content-filter-groundedness"&gt;check OpenAI’s homework&lt;/a&gt; if you want. &lt;a class=footnote-backref href=#fnref:3:serious-ai-product-2026-9 title="Jump back to footnote 3 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:4:serious-ai-product-2026-9&gt;
&lt;p id=fn:4&gt;Given the relatively slow speed of the justice system and the mainstream
press around the world, we probably will not hear about whether people are
managing to incidentally break through these guard rails to self harm right
now, but there are no shortage of stories still being &lt;em&gt;reported&lt;/em&gt; right now
where people were still doing just that, such as in &lt;a href="https://www.npr.org/2026/08/18/nx-s1-5929575/ai-suicide-risks-mental-health"&gt;this
story&lt;/a&gt;
where the effect of the much vaunted “guard rails” in 2025 was that if you
wanted it to write you a suicide note, it would refuse twice but acquiesce
on the third try.  I don’t see any reason to believe this fundamental issue
has been addressed in the meanwhile, since it had been happening for years
at that point. &lt;a class=footnote-backref href=#fnref:4:serious-ai-product-2026-9 title="Jump back to footnote 4 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:5:serious-ai-product-2026-9&gt;
&lt;p id=fn:5&gt;In this tutorial we can also see an incredibly rosy scenario presented,
where a long-running workflow effortlessly compresses all of the necessary
information into the new context, even if it uses a lower-fidelity model to
do so, rather than the tangled and gnarly problem of problems which really
&lt;em&gt;are&lt;/em&gt; too big to fit in the context, which is to say, “most real-world
problems”.  This presents the context limit instead as a minor speedbump to
be worked around rather than the fundamental flaw in LLM tooling. &lt;a class=footnote-backref href=#fnref:5:serious-ai-product-2026-9 title="Jump back to footnote 5 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:6:serious-ai-product-2026-9&gt;
&lt;p id=fn:6&gt;A heavily fictionalized seaport.  This is not how actual dockworkers
work.  This is a simplistic metaphor about incidental benefits of
instrumental tasks, it is not supposed to delve deeply into the mechanics
of maritime shipping. In particular I know that cranes are more reliable
than this and this is not actually how you would respond to a crane
malfunction anyway.  Feel free to share fun facts about maritime shipping
if that is your special interest but please do not @ me to &lt;em&gt;correct&lt;/em&gt; this
metaphor. &lt;a class=footnote-backref href=#fnref:6:serious-ai-product-2026-9 title="Jump back to footnote 6 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:7:serious-ai-product-2026-9&gt;
&lt;p id=fn:7&gt;To my knowledge, none of their publicly-traded employers have posted a
measurable improvement to efficiency outside the margin of error. &lt;a class=footnote-backref href=#fnref:7:serious-ai-product-2026-9 title="Jump back to footnote 7 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:8:serious-ai-product-2026-9&gt;
&lt;p id=fn:8&gt;Or any similar methodology.  I don’t need people to adopt the exact
practice that I proposed there. &lt;a class=footnote-backref href=#fnref:8:serious-ai-product-2026-9 title="Jump back to footnote 8 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;&lt;/body&gt;</content><category term="misc"></category><category term="ai"></category><category term="llm"></category></entry><entry><title>Who Is Open Source About?</title><link href="https://blog.glyph.im/2026/09/who-is-open-source-about.html" rel="alternate"></link><published>2026-09-24T17:50:00-07:00</published><updated>2026-09-24T17:50:00-07:00</updated><author><name>Glyph</name></author><id>tag:blog.glyph.im,2026-09-24:/2026/09/who-is-open-source-about.html</id><summary type="html">&lt;p&gt;Open source comprises a complex social web of ongoing relationships,
not a simple one-way “gift” of value from maintainer to user.&lt;/p&gt;</summary><content type="html">&lt;body&gt;&lt;p&gt;Open source is, at least in part, &lt;strong&gt;about &lt;em&gt;you&lt;/em&gt;&lt;/strong&gt;, where “you” refers to the
user.&lt;/p&gt;
&lt;h2 id=open-source-is-not-about-open-source-is-not-about-you&gt;Open Source Is Not About “Open Source Is Not About You”&lt;/h2&gt;
&lt;p&gt;In other words: &lt;a href="https://gist.github.com/richhickey/1563cddea1002958f96e7ba9519972d9"&gt;Rich Hickey was wrong when he wrote “Open Source Is Not About
You”&lt;/a&gt; and
I’m tired of pretending otherwise.&lt;/p&gt;
&lt;p&gt;Of course he’s not &lt;em&gt;completely&lt;/em&gt; wrong, or his famous post would not have
resonated quite so much in the first place.  Obnoxious users who demand their
&lt;em&gt;personal&lt;/em&gt; use-cases be immediately addressed by volunteer maintainers for free
should indeed be viewed as the pariahs that they are.  Similarly, corporate
users who want free support from the community that supplies their
infrastructure to lower their costs.  As should those who
&lt;a href="https://alexgaynor.net/2015/mar/30/red-hat-open-source-community/"&gt;profit&lt;/a&gt;
from this type of externalization by their own customers.&lt;/p&gt;
&lt;p&gt;But the exchange of “open source” (or even “free software”) is not as simple as
“I have prepared some software for you, please enjoy it, you have no right to
complain”, and maintainers ought to have a precise understanding of the costs
and benefits — as well as the ethical implications — of that exchange.&lt;/p&gt;
&lt;p&gt;Right now we barely even articulate that the exchange &lt;em&gt;exists&lt;/em&gt;, let alone that
it establishes a long-term, subtle, and implicit relationship between
maintainer and user.&lt;/p&gt;
&lt;p&gt;Let’s fix that.&lt;/p&gt;
&lt;div class=aside&gt;
&lt;h3 id=a-brief-aside-about-meta-ethics&gt;A Brief Aside about Meta-Ethics&lt;/h3&gt;
&lt;p&gt;When we talk about “obligations” and “rights”, of “shoulds” and “musts”, we are
constructing an ethical system. The &lt;em&gt;purpose&lt;/em&gt; of such a system is to develop
social expectations and social consequences.  There is not much use in me
telling you that you are &lt;em&gt;transcendentally evil&lt;/em&gt; for failing to follow some
arbitrary recommendation that I have.  But I am implying that I believe there
should be consequences for your behavior.  I am also implying that there
probably already &lt;em&gt;are&lt;/em&gt; some consequences, and they’re just not written down
anywhere yet.&lt;/p&gt;
&lt;p&gt;Therefore, a post like this, where I say that we &lt;em&gt;should&lt;/em&gt; view our social
obligations in a certain way, that is the &lt;em&gt;beginning&lt;/em&gt; of a broader social
conversation.  I think there should be some consequences, so I am gesturing
towards that possibility.  Exactly what consequences?&lt;/p&gt;
&lt;p&gt;For now, I’m not sure.  Let’s figure it out.&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id=what-are-we-doing-when-we-do-an-open-source&gt;What Are We Doing When We Do An Open Source?&lt;/h2&gt;
&lt;p&gt;Hickey, and his many acolytes in the years since his fateful post, asserts that
the process of “open source” goes like this:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Maintainer makes a thing, and makes it available to users as a gift.&lt;ol&gt;
&lt;li&gt;Maintainer may “love working with the team”.&lt;/li&gt;
&lt;li&gt;Maintainer may be “proud of the work we do”.&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;Users accept the gift, and extract utility from it.&lt;ol&gt;
&lt;li&gt;(Users MUST be grateful for this.)&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;A tiny fraction of users reciprocally contribute to the thing.&lt;ol&gt;
&lt;li&gt;(Maintainers may be grateful for this.)&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;He makes various oblique references to the specific activities of his company,
which does things vaguely related to his projects for money&lt;sup id=fnref:1:who-is-open-source-about-2026-9&gt;&lt;a class=footnote-ref href=#fn:1:who-is-open-source-about-2026-9 id=fnref:1&gt;1&lt;/a&gt;&lt;/sup&gt;.  These
activities are exclusively characterized as for “customers”, however, a subset
of the aforementioned users so tiny (“fewer than 1%”) as to nearly be an
entirely distinct group.&lt;/p&gt;
&lt;p&gt;Breezing past this process in an essay about obnoxious users demanding things
they are not entitled to, one might nod along, as this sounds mostly sensible.
Giving gifts is nice.  I too love working with good teams and taking pride in
things.&lt;/p&gt;
&lt;p&gt;Examined more closely, however, it starts to logically fall apart.  If you have
consulting clients and that’s where all of your money is coming from, why are
you &lt;em&gt;bothering&lt;/em&gt; (as he repeatedly insists) “doing [things] for the community”?
What was the point of releasing this code in the first place?  You could love
working with your team and be proud of the work that you do in a lot of
different contexts; why bother implicating this horde of entitled and obnoxious
people, if that’s all you’re getting out of it?  &lt;em&gt;What’s in it for you?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;If we’ve left out something as fundamental as “why is the maintainer doing
this”, perhaps this story leaves out some &lt;em&gt;other&lt;/em&gt; important bits as well.&lt;/p&gt;
&lt;h2 id=why-are-you-doing-this&gt;Why Are You Doing This?&lt;/h2&gt;
&lt;p&gt;There are &lt;em&gt;many&lt;/em&gt; possible motivations for releasing and maintaining open source
software.  They are often subtle, often overlapping, and rarely clearly stated.
Maintainers are not a monolith and not everyone does it for similar reasons.
But let’s review a few reasons that someone might want to contribute.&lt;/p&gt;
&lt;h3 id=reputation&gt;Reputation&lt;/h3&gt;
&lt;p&gt;One reason that you might want to release some open source software is
&lt;em&gt;advertising&lt;/em&gt;.  The most common form of this is self-promotion; if you are a
visible, prominent contributor to an open source project, it stands to reason
that you will have an easier time finding work in the domain of that project.&lt;/p&gt;
&lt;p&gt;If you operate a consultancy, as Rich Hickey did at the time of his famous
rant, then this reputational currency translates into advertising for your
services.  It’s a practical demonstration of the skills of your team.&lt;/p&gt;
&lt;p&gt;The trade in this benefit is most like the traditional “gift economy” that open
source has been compared to.  You give the code to your users, which has some
value, but the users give you back some reputation, in the form of their
attention, their esteem, and possibly even their money if they become customers
or employers.&lt;/p&gt;
&lt;h3 id=influence&gt;Influence&lt;/h3&gt;
&lt;p&gt;Infrastructure is the most popular type of open source for a good reason.
Programmers working on a problem are often hemmed in by sclerotic architectural
choices which prevent them from solving problems in the way that they’d prefer
to solve them.  Major infrastructural investments are difficult to justify in a
planning process, as their benefits are hard to prove.  Sometimes the benefits
are highly personal; different engineers have different aesthetic preferences
about what types of equally-valid solutions they’d prefer to work with.&lt;/p&gt;
&lt;p&gt;If you can develop your preferred type of solution and release it as open
source, then you can influence &lt;em&gt;how&lt;/em&gt; everyone else solves this type of problem.
As an individual, such a position of influence can allow you to have some
transferable expertise between employers. You know how to use the tool you
developed, so you can be very quick and effective with it, and you can shape it
to your ongoing taste over time.&lt;/p&gt;
&lt;p&gt;If you’re an employer, and you can get everyone &lt;em&gt;else&lt;/em&gt; to use your open source
thing&lt;sup id=fnref:2:who-is-open-source-about-2026-9&gt;&lt;a class=footnote-ref href=#fn:2:who-is-open-source-about-2026-9 id=fnref:2&gt;2&lt;/a&gt;&lt;/sup&gt;, this can reduce &lt;em&gt;both&lt;/em&gt; your hiring and training costs.  Potential
employees can read the code, see that it’s good, and want to work at a place
that produces good code like that.  They can also read the code and become
familiar with it in advance of coming to work for you, which means that you
have a ready supply of developers who already know how your internal systems work.&lt;/p&gt;
&lt;p&gt;The trade in this benefit is more like “soft power” than a gift economy.  You
give the code to your users, which has some value, but the users give you back
the ability to dictate their technological agenda.  You gain both the ability
to influence their initial direction, and, as part of ongoing maintenance, to
dictate their behavior over time.&lt;/p&gt;
&lt;h3 id=improvement&gt;Improvement&lt;/h3&gt;
&lt;p&gt;As an engineer, you might want to improve your own skills.  Writing something
proprietary and commercial cuts against this in two ways.&lt;/p&gt;
&lt;p&gt;First, you will want to build something that already exists within your skill
set, so that it will attract commercial interest and actually be competitive.
Within the context of a larger team, you will want to personally be able to be
immediately effective for similar reasons.  But you still need a way to learn
new things.&lt;/p&gt;
&lt;p&gt;Second, you will want to build something somewhat secretively, so that the
value you are producing is captured rather than released to the community.
This means that you will be cut off from external sources of expert feedback.&lt;/p&gt;
&lt;p&gt;As an organization, you might want to build the skills of your staff in similar ways.&lt;/p&gt;
&lt;p&gt;The trade in this benefit is code for knowledge.  You release the code or
changes, and in return you expect your users to provide you good bug reports,
and to induce at least some of them to become co-developers.&lt;/p&gt;
&lt;h3 id=outsourcing&gt;Outsourcing&lt;/h3&gt;
&lt;p&gt;As an engineer, you can only do so much on your own.  Perhaps you want to have
some influence over your infrastructure so you want to write it, but you also
want to have a communal place to keep your infrastructure such that you can
&lt;em&gt;make&lt;/em&gt; a change to something to suit your needs, but you know that even if you
walk away, someone else will &lt;em&gt;maintain&lt;/em&gt; that change and keep it working across years or even decades of changes to underlying platforms, hardware, etc.&lt;/p&gt;
&lt;p&gt;This sort of communal maintenance effort can be shared among all interested
participants; if a thousand companies all need the same tool, if even a few
dozen can share it, that reduces even their &lt;em&gt;own&lt;/em&gt; load massively, let alone
everyone else’s.&lt;/p&gt;
&lt;p&gt;The trade in this benefit is more complex, since there’s less symmetry between
the main maintainer and peripheral community members who also contribute code.
The main maintainer is actually trading a &lt;em&gt;namespace&lt;/em&gt;, a central place for
people to contribute, coordinate, and release changes, rather than the &lt;em&gt;code&lt;/em&gt;.
They are a sort of market maker where then all the other contributors trade
code &lt;em&gt;for&lt;/em&gt; code within that market-ish structure.&lt;/p&gt;
&lt;p&gt;In practice, this motivation produces a &lt;a href="https://en.wikipedia.org/wiki/Volunteer's_dilemma"&gt;game theory
problem&lt;/a&gt; where, when
maintenance drops below a critical threshold, it creates a &lt;a href="https://en.wikipedia.org/wiki/Heartbleed"&gt;big enough
crisis&lt;/a&gt; that at least &lt;em&gt;some&lt;/em&gt;
freeloading stakeholders will be forced to start making contributions.&lt;/p&gt;
&lt;p&gt;Ultimately, however, this saves &lt;em&gt;all&lt;/em&gt; involved parties a ton on maintenance,
more eager volunteers who do not freeload in the first place get all the other
benefits mentioned above as well.&lt;/p&gt;
&lt;div class=aside&gt;
&lt;h3 id=a-brief-aside-about-your-chart-of-accounts&gt;A Brief Aside about your Chart of Accounts&lt;/h3&gt;
&lt;p&gt;Most companies account for open source maintenance work as simple overhead on
ongoing projects.  Sometimes it’s CapEx, sometimes it’s OpEx, but it’s just
“whoever happens to be working on this thing to support whatever random product
it’s a part of”.&lt;/p&gt;
&lt;p&gt;This type of accounting creates distorting incentives, because it doesn’t
recognize all the benefits above.  Under such a fiscal regime, ongoing healthy
maintenance becomes a
&lt;a href="https://www.businessinsider.com/zirp-end-of-cushy-big-tech-job-perks-mass-layoffs-2024-2"&gt;ZIRP&lt;/a&gt;
because when resources are more constrained, this apparent indulgence gets
corrected.&lt;/p&gt;
&lt;p&gt;The ancillary benefits that open source creates ought to be properly
recognized.  It shouldn’t just be buried as Wages or IT or whatever.  If it’s
helping you hire better engineers, some of that expense should be allocated to
Recruitment Costs.  If it’s materially improving your reputation among your
customer base, some of it should go to Goodwill.  If it’s getting your product
in front of developers who are your customers, it should be in Marketing.  Most
importantly, if maintenance on an open source project is actually helping you
maintain your enterprise-wide platform, it should not be squirreled away in
some small team who happened to be the first one to adopt it.&lt;sup id=fnref:3:who-is-open-source-about-2026-9&gt;&lt;a class=footnote-ref href=#fn:3:who-is-open-source-about-2026-9 id=fnref:3&gt;3&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;Exactly how these costs should be allocated and cross-charged to different
departments depends heavily upon your organization and your specific chart of
accounts.  But “whatever, it’s just part of the software product” or “I guess
it’s DevRel because the SDK is in there” is guaranteed to have your open source
organization destroyed along with all those side-benefits the next time that
there’s a &lt;a href="https://cepr.net/publications/ai-bubble-monitor/"&gt;cash crunch&lt;/a&gt;.&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id=the-things-that-arent-supposed-to-be-benefits&gt;The Things that Aren’t Supposed To Be Benefits&lt;/h2&gt;
&lt;p&gt;These categories &lt;em&gt;could&lt;/em&gt; be made as explicit, rational trade-offs, even if they
are often implicit and subtle in practice.  They are transactions where the
maintainer gets something and the user gets something.&lt;/p&gt;
&lt;p&gt;However, not everything that you are getting as a maintainer is something you
are actually &lt;em&gt;supposed to use&lt;/em&gt; to your own benefit.  Being given trust in
service of a responsibility is not a transaction.&lt;/p&gt;
&lt;h3 id=oops-all-root-shells&gt;“Oops, All Root Shells”&lt;/h3&gt;
&lt;p&gt;Open source code &lt;em&gt;is code&lt;/em&gt;.  In our modern world of &lt;a href="https://pyvideo.org/pybay-2024/when-arbitrary-code-execution-is-working-as-intended-what-code-is-python-supposed-to-execute.html"&gt;absolutely pathetic
sandboxing&lt;/a&gt;,
installing code from somebody else gives them control over your system, even if
it is somewhat indirect.&lt;/p&gt;
&lt;p&gt;There is an unwritten rule that if I create an open source library, and you use
it, it probably &lt;em&gt;shouldn’t&lt;/em&gt; have a backdoor in it that gives me the credentials
to your bank account.  There is a trust relationship between the user and the
maintainer, and here, we see the first &lt;em&gt;obligation&lt;/em&gt; that the maintainer has.
The maintainer is obligated &lt;em&gt;not to use the user’s computer for their own
gain&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;This rule might seem obvious and straightforward.  It might even seem unfair to
you that I call the rule “unwritten”, because the rule &lt;em&gt;is&lt;/em&gt;, in fact, written
down in a few places: for example, in the &lt;a href="https://docs.npmjs.com/policies/open-source-terms#acceptable-content"&gt;npm Acceptable Content
Policy&lt;/a&gt;,
it says right there:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;A few examples of unacceptable content:&lt;/p&gt;
&lt;p&gt;…&lt;/p&gt;
&lt;ol start=3&gt;
&lt;li&gt;Content containing malicious computer code, such as computer viruses,
   computer worms, rootkits, back doors, or spyware. This includes content
   submitted for research purposes. Tools designed and documented explicitly
   to assist in security research are acceptable, but exploits and malware
   that use the npm registry as a deployment or delivery vector are not.&lt;/li&gt;
&lt;/ol&gt;
&lt;/blockquote&gt;
&lt;p&gt;I think we can all agree that a script which steals your bank credentials and
sends them to me to buy a totally sick jet ski would qualify as “malware”, so
clearly that is forbidden.&lt;/p&gt;
&lt;p&gt;There is also an enormous gray area here.  npm also explicitly &lt;em&gt;allows&lt;/em&gt;
“Information on how to pay, donate to, and otherwise support Package
development”, but then goes on to explicitly &lt;em&gt;forbid&lt;/em&gt; “Packages that display
ads at runtime, on installation, or at other stages of the software development
lifecycle, such as via npm scripts.”&lt;sup id=fnref:4:who-is-open-source-about-2026-9&gt;&lt;a class=footnote-ref href=#fn:4:who-is-open-source-about-2026-9 id=fnref:4&gt;4&lt;/a&gt;&lt;/sup&gt; How are the lines drawn around these
gray areas?  “npm will continue to apply its judgment when deciding what
content is acceptable.”&lt;/p&gt;
&lt;p&gt;But also... this is forbidden &lt;em&gt;by npm&lt;/em&gt;, not by the transcendental nature of
“open source”.  I could give away code that displays all kinds of ads to its
users as a “gift” on my website.  The exact structure of this policy is not
uncommon, but it also isn’t exactly the same as other such sites.  &lt;a href="https://policies.python.org/pypi.org/Acceptable-Use-Policy/"&gt;PyPI, for
example&lt;/a&gt;,
explicitly bans “cryptocurrency mining”, which NPM does not.  Is cryptocurrency
mining “not open source”?  A lot of judgement calls are happening here about
what is allowable in these “gifts” that you are giving to your users.&lt;/p&gt;
&lt;p&gt;But I digress.&lt;/p&gt;
&lt;p&gt;My point is that policy-making around this concept is not clear, there are lots
of little disagreements around the edges, but there is a very strong consensus
that while the user is giving you their trust here, that is &lt;strong&gt;not a trade&lt;/strong&gt;.
The deal is not “you give the user some code, the user gives you unlimited
compute and access to all their financial accounts”.  The user has made
themselves vulnerable to your code on the strength of your reputation.&lt;/p&gt;
&lt;p&gt;This creates an obligation for you to not do anything evil with that code,
either intentionally or through negligence.&lt;/p&gt;
&lt;h4 id=security-updates-are-just-command-and-control-in-a-funny-hat&gt;Security Updates Are Just Command And Control In A Funny Hat&lt;/h4&gt;
&lt;p&gt;All of this is just about the initial download of some code, and that is the
way that Rich Hickey describes it, as if you just grabbed some code off a web
page and put it in a folder that you like on your desktop.  But that is not how
open source relationships work today, if indeed it ever was.&lt;/p&gt;
&lt;p&gt;The way it works today is that you add a dependency to your &lt;code&gt;pyproject.toml&lt;/code&gt; or
your &lt;code&gt;package.json&lt;/code&gt; or your &lt;code&gt;Cargo.toml&lt;/code&gt; and now your users are vulnerable not
just to whatever you happened to upload in the first place, but to &lt;em&gt;whoever
happens to have your package index credentials&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;This creates an obligation to maintain an operational security posture that
protects your users from malicious updates.&lt;/p&gt;
&lt;h3 id=the-roadmap-is-someones-life&gt;The Roadmap Is Someone’s Life&lt;/h3&gt;
&lt;p&gt;Another kind of trust that the user is placing in you is the trust that you are
going to have at least &lt;em&gt;some&lt;/em&gt; kind of regard for their usage of your software.&lt;/p&gt;
&lt;p&gt;In a perfect world, the user’s expectations could be clearly circumscribed.
Whatever ongoing maintenance you commit to perform would be encapsulated in
clear policies that you’d write up in advance, about exactly what kind of
security response policy you have, how you will communicate when you no longer
have the resources for maintenance, and so on.&lt;/p&gt;
&lt;p&gt;But anyone who has been involved in any project at anything but the most
extreme tier of operational maturity knows that 99% of the ecosystem relies on
a set of loose conventions around how all that stuff works.  We expect that
maintainers will generally be around, that they’ll use existing tools like an
issue tracker for triaging user bugs, GHSA and CVEs for security reporting,
that they will mark the project as “archived” and maybe do a final release
before abandoning it, that they will maintain a ChangeLog explaining at least a
little bit of what is going on.&lt;/p&gt;
&lt;p&gt;Users assume that those conventions will be followed when there are any gaps in
explicit policy, or indeed if policy is lacking entirely.  This assumption is
reasonable, because otherwise nobody could ever &lt;em&gt;use&lt;/em&gt; any open source without a
stack of service contracts that nobody has any time to write.&lt;/p&gt;
&lt;p&gt;The strongest such convention is that an actively maintained program will, at
least, more or less &lt;em&gt;keep doing what it does&lt;/em&gt; as time goes on.  A user who has
elected to use a bit of open source software has made themselves vulnerable to
changes and breakages in that software by the mere fact of using it.  In the
time that they have used it and invested in it, they have &lt;em&gt;not&lt;/em&gt; invested in:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;creating&lt;/em&gt; alternative software to meet their needs,&lt;/li&gt;
&lt;li&gt;&lt;em&gt;maintaining data&lt;/em&gt; in formats that other software can read, or&lt;/li&gt;
&lt;li&gt;&lt;em&gt;learning how to use&lt;/em&gt; existing alternative software.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This can, and does, go badly wrong, when those expectations are mismatched.&lt;/p&gt;
&lt;h4 id=how-it-goes-wrong&gt;How It Goes Wrong&lt;/h4&gt;
&lt;p&gt;Let’s say a maintainer creates an open source paint program, OpenPaint.&lt;/p&gt;
&lt;p&gt;An artist, known for their unique style of making blended collages, switches
from their previous app, ProprietaryPaint, to this new OpenPaint to make these
culturally significant works of art.  However, the maintainer decides that the
‘blend’ tool is kind of a pain to maintain, and they remove it in OpenPaint 2.&lt;/p&gt;
&lt;p&gt;A few months later, the artist’s operating system vendor issues a security
update that breaks OpenPaint, because older versions of OpenPaint were
unknowingly abusing some platform API.&lt;/p&gt;
&lt;p&gt;The maintainer releases a new OpenPaint 2.0.1 that addresses this
incompatibility, but doesn’t care about version 1.x any more so they don’t
bother to update that one.&lt;/p&gt;
&lt;p&gt;This places the artist in an impossible situation.  They can stay on an old
version of their operating system, putting all their personal data at risk.  Or
they can upgrade to the new operating system, effectively either cutting off
access to their livelihood, or forcing them to change their art style entirely.&lt;/p&gt;
&lt;p&gt;Now, proprietary software can place users in similarly untenable positions (and
in fact, it is &lt;em&gt;more often&lt;/em&gt; proprietary software that does).  But does the
openness completely remove &lt;em&gt;any&lt;/em&gt; obligation for this consideration?  Should
the OpenPaint team have to at least &lt;em&gt;communicate&lt;/em&gt; the reasons for doing this,
to give the artist some recourse?&lt;sup id=fnref:5:who-is-open-source-about-2026-9&gt;&lt;a class=footnote-ref href=#fn:5:who-is-open-source-about-2026-9 id=fnref:5&gt;5&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;The only thing that “open source” does is that it allows the artist to pay a
prohibitive amount of money to a new maintenance team to create a fork.  This
is rarely the kind of thing that individuals can manage.&lt;/p&gt;
&lt;p&gt;This creates an obligation to &lt;em&gt;at least consider how your users might be
relying on you&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;This is the most complex obligation of the bunch.  Obviously it does not
entitle every single user to infinite work from the maintainer, but it also
shouldn’t entitle the user to &lt;em&gt;nothing&lt;/em&gt; for having trusted these subtle implied
claims that the maintainer is making by making their work public.&lt;/p&gt;
&lt;p&gt;It is a nuanced and ongoing negotiation and I do not think we have a clear
moral intuition about how it should work out.  But we do need to figure out a
way to work it out.&lt;/p&gt;
&lt;p&gt;It also raises a clarifying question.&lt;/p&gt;
&lt;h2 id=why-are-we-even-doing-this-and-who-are-we-doing-it-for&gt;Why Are We Even Doing This, and Who Are We Doing It For?&lt;/h2&gt;
&lt;p&gt;People generally like to do things for more than one reason.  We live in an
economy where people need to make money, but we mostly prefer to make that
money doing things that are useful, and that make other people happy.&lt;/p&gt;
&lt;p&gt;So, yes, we create open source for self-interested reasons to improve our
reputations, to improve our skills, to increase our influence and to share our
maintenance burdens.  In so doing we take on &lt;em&gt;some&lt;/em&gt; level of obligation to not
abuse the trust that is placed in us, even if that level of obligation is not
clear.&lt;/p&gt;
&lt;p&gt;But if we are not doing it to &lt;em&gt;serve&lt;/em&gt; those users at least a &lt;em&gt;little&lt;/em&gt; bit, then
those motivations are going to quickly ring hollow.  We will not increase our
reputation with a person if we respond to their every request by telling them
that we owe them nothing and that their opinions are worthless.  We will not
gain influence over a community if we ignore their desires.&lt;/p&gt;
&lt;p&gt;Many interactions with open source maintainers are unnecessarily adversarial.
This is of course partially the fault of those users, who should calibrate
their expectations appropriately.&lt;/p&gt;
&lt;p&gt;Still: maintainers could do a better job of listening &lt;em&gt;before&lt;/em&gt; these
interactions become toxic.  There’s no reason that “open source users” should
be an especially toxic group of people.  At this point in history, that group
is basically just … people with computers.&lt;/p&gt;
&lt;p&gt;It’s like that old truism.  If you meet one person who is a jerk to you, that’s
their problem.  But if everyone you meet, everywhere you go, is constantly
abrasive to you and treats you like you’re doing something wrong, maybe it’s
time to look inward.&lt;/p&gt;
&lt;p&gt;If all open source users are entitled assholes, maybe it’s time to look for a
structural problem.&lt;/p&gt;
&lt;h3 id=surprise-its-about-ai-again&gt;Surprise, It’s About AI Again&lt;/h3&gt;
&lt;p&gt;Sigh.&lt;sup id=fnref:6:who-is-open-source-about-2026-9&gt;&lt;a class=footnote-ref href=#fn:6:who-is-open-source-about-2026-9 id=fnref:6&gt;6&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Users &lt;/em&gt;&lt;em&gt;hate&lt;/em&gt;&lt;em&gt; slop.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;I know, dear AI-positive reader, &lt;em&gt;your&lt;/em&gt; AI outputs are different from everyone
else’s, &lt;em&gt;you&lt;/em&gt; aren’t pushing thoughtless slop into your code, just because
everyone else is and it is the inevitable terminus of using those tools.  You
aren’t “lazy vibe coding” with Claude, you’re doing “responsible agentic
engineering”, which is different because you’re just built different.&lt;/p&gt;
&lt;p&gt;Still, humor me, for a moment.  Your &lt;em&gt;users&lt;/em&gt; don’t know that.  They know what
it looks like when products that they like adopt slop.  They know that they
will start &lt;a href="https://www.theregister.com/2026/02/26/veracode_security_ai/"&gt;leaking
data&lt;/a&gt;.
Developers know that it will make them &lt;a href="https://www.darkreading.com/application-security/flaws-claude-code-developer-machines-risk"&gt;personally less
secure&lt;/a&gt;.
They know that they can expect &lt;a href="https://www.pcmag.com/news/vibe-coding-fiasco-replite-ai-agent-goes-rogue-deletes-company-database"&gt;more
outages&lt;/a&gt;
and that your code will &lt;a href="https://arstechnica.com/ai/2026/03/after-outages-amazon-to-make-senior-engineers-sign-off-on-ai-assisted-changes/"&gt;inexorably decline in
quality&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;In other words, your users are going to assume that this means you are
violating that final obligation that the software should &lt;em&gt;keep working&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;Your users are going to tell you to stop, and they are probably going to get
mad.  Maybe you, or a plurality of your team, &lt;em&gt;also&lt;/em&gt; want to stop, maybe you
disagree with them, but in any case you need some way to &lt;em&gt;have that
conversation&lt;/em&gt; in a way that does not immediately overflow into every adjacent
discussion forum.  Users need to feel welcome in some space so they can have
the discussion &lt;em&gt;in&lt;/em&gt; that space, and not explode out into a thousand different
group chats and social media threads.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;This&lt;/em&gt; post was inspired by yet another prominent open source community
discourse where a ton of angry users showed up to yell at developers to stop
accepting LLM-generated code.  I’m not going to link to any of these, because
we don’t need any more fuel for the discourse fire.  But there is more than one
such case and the pattern is becoming familiar.&lt;/p&gt;
&lt;p&gt;On social media - usually BlueSky or Mastodon, but sometimes a user group
forum - users become aware of some AI-adjacent policy.  They show up in a horde
to the developer forum or mailing list.  They loudly start demanding the
project take a hard stand&lt;sup id=fnref:7:who-is-open-source-about-2026-9&gt;&lt;a class=footnote-ref href=#fn:7:who-is-open-source-about-2026-9 id=fnref:7&gt;7&lt;/a&gt;&lt;/sup&gt; against AI. This pressure is simultaneous, but
uncoordinated; extremely repetitive, very diverse, often inconsistent, and
pretty stressful, especially if you’re a burnt-out maintainer with other things
to be doing who may not even like AI yourself in the first place.&lt;/p&gt;
&lt;p&gt;Believe me, I get it.  It can be very unpleasant to deal with.&lt;/p&gt;
&lt;p&gt;Like most problems that AI is causing, though, it’s not really an “AI” problem
as much as it is a pre-existing dumpster fire that “AI” is pouring gasoline
onto.  In this case, an online mob is the language of the unheard&lt;sup id=fnref:8:who-is-open-source-about-2026-9&gt;&lt;a class=footnote-ref href=#fn:8:who-is-open-source-about-2026-9 id=fnref:8&gt;8&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;h2 id=if-users-are-mad-its-probably-already-too-late-but-maybe-you-can-get-ready-for-next-time&gt;If Users Are Mad It’s Probably Already Too Late (But Maybe You Can Get Ready For Next Time)&lt;/h2&gt;
&lt;p&gt;One day, all of a sudden, you’re getting feedback from a bunch of users that
are using inappropriate channels to complain. But did they already have
&lt;em&gt;appropriate&lt;/em&gt; channels to use?&lt;/p&gt;
&lt;p&gt;Did you have a place for people to congregate and discuss your project?  To
make orderly complaints in a way that will be legible to you?  Or do you just
have a GitHub Issues page, which non-technical users have no idea how to
interact with, and a forum for developers, where users don’t know the norms and
any arriving brigade of pissed-off users will be seen as disruptive and
inappropriate?&lt;/p&gt;
&lt;p&gt;I don’t want to be throwing any stones from within my particular glass house.
Setting up such a place has gotten harder over the years.  I don’t really have
one, either.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Could&lt;/em&gt; I have one, though?  IRC has been slowly dying, mailing lists are
unpopular and present increasingly annoying moderation challenges, forum
software is expensive to operate and keep maintained, Discord is a confusing
mess and the upshot of all of this is every community needs community
management and forum moderation.  Which means that for my own small solo
projects, I couldn’t possibly have such infrastructure because such
infrastructure requires a &lt;em&gt;dedicated second person&lt;/em&gt; to maintain it, and until
someone volunteers for that, it’s not really feasible.  Even for my &lt;a href="https://twisted.org"&gt;larger
projects&lt;/a&gt; you’d be surprised how slim of a skeleton crew
we are getting by with, and we definitely don’t have a whole spare maintainer
to go manage this, especially as we are under attack from the slopocalypse
ourselves.&lt;/p&gt;
&lt;p&gt;The nature of open source community is that most communities start too small to
need such a thing, grow incrementally until one day they are suddenly &lt;em&gt;way&lt;/em&gt; too
big and needed one yesterday, and then suddenly they are too small again when
interest wanes even a little bit.  Even as we need it more and more, building
and maintaining community infrastructure remains a challenge.&lt;/p&gt;
&lt;p&gt;Even so, having a dedicated place for &lt;em&gt;users&lt;/em&gt; — not maintainers — to converse
amongst themselves, be an actual community, and present feedback to the
developers, is fast becoming a necessary component of a successful community
and not a nice-to-have.&lt;/p&gt;
&lt;h2 id=in-conclusion&gt;In Conclusion&lt;/h2&gt;
&lt;p&gt;As trying as it can be sometimes, we maintainers all do get something out of
open source, and it is good to be honest with your users — and with yourself —
exactly &lt;em&gt;what&lt;/em&gt; you want to get out of it.  In order to know whether the &lt;a href="https://en.wiktionary.org/wiki/the_juice_is_worth_the_squeeze"&gt;juice
is worth the
squeeze&lt;/a&gt;, we
must know both what the juice is, &lt;em&gt;and&lt;/em&gt; what the squeeze is.&lt;/p&gt;
&lt;p&gt;Part of the metaphorical squeeze &lt;em&gt;is&lt;/em&gt; a set of obligations, and those are the
most poorly defined of all.  We should try to be clear about what those are
too.  Both about exactly what we believe we are signing up for, and also, about
how we are willing to let our users hold us to account for them.  Codes of
conduct are a start here, but only the absolute barest bare minimum; “do not
harass your colleagues or your users” is not a standard of excellence to aspire
to, it’s just basic manners.&lt;/p&gt;
&lt;p&gt;I can’t tell you exactly what your obligations are, only try to gesture at my
idea of the outlines of the fuzzy moral intuition we’ve all been implicitly
sharing up until now.&lt;/p&gt;
&lt;p&gt;Drawing this line is not just for the benefit of the users, either.
Maintainers already feel pressure, we already feel obligations.  We &lt;em&gt;resent&lt;/em&gt;
that feeling of obligation. While there are a diverse array of reasons for that
resentment, one big one is that it’s not clear, even to ourselves &lt;em&gt;where the
obligations end&lt;/em&gt;.  Lashing out by saying “I promised nothing and I owe you
nothing!” followed by some choice expletives feels cathartic, but it doesn’t
really solve the problem, because we clearly don’t really believe that’s where
the line is, or we would have already stopped there.  We wouldn’t feel the need
to say it.&lt;/p&gt;
&lt;p&gt;It is going to be a very big collective endeavor to figure out exactly where
that line is. The best time to have gotten started on that endeavor was 50
years ago.&lt;/p&gt;
&lt;p&gt;But the second best time is today.&lt;/p&gt;
&lt;h2 id=acknowledgments&gt;Acknowledgments&lt;/h2&gt;
&lt;p class=update-note&gt;Thank you to &lt;a href="/pages/patrons.html"&gt;my patrons&lt;/a&gt; who are supporting my writing on
this blog.  If you like what you’ve read here and you’d like to read more of
it, or you’d like to support my &lt;a href="https://github.com/glyph/"&gt;various open-source
endeavors&lt;/a&gt;, you can &lt;a href="/pages/patrons.html"&gt;support my work as a
sponsor&lt;/a&gt;!&lt;sup id=fnref:9:who-is-open-source-about-2026-9&gt;&lt;a class=footnote-ref href=#fn:9:who-is-open-source-about-2026-9 id=fnref:9&gt;9&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;div class=footnote&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=fn:1:who-is-open-source-about-2026-9&gt;
&lt;p id=fn:1&gt;Somewhat to everyone’s surprise, I, too, do things for money, like
writing this post. Please remember to &lt;a href="https://blog.glyph.im/pages/patrons.html"&gt;like and
subscribe&lt;/a&gt; &lt;a class=footnote-backref href=#fnref:1:who-is-open-source-about-2026-9 title="Jump back to footnote 1 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:2:who-is-open-source-about-2026-9&gt;
&lt;p id=fn:2&gt;Whether it was originally yours, or developed by an employee who happened
to be on staff at the time, or adopted by an employee who just started
contributing to it a lot, in any of these scenarios a company can benefit
from increased consistency and increased familiarity. &lt;a class=footnote-backref href=#fnref:2:who-is-open-source-about-2026-9 title="Jump back to footnote 2 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:3:who-is-open-source-about-2026-9&gt;
&lt;p id=fn:3&gt;If the rule is that they must forever endure the searing budgetary pain
of gripping the &lt;a href="https://en.wikipedia.org/wiki/Hot_potato"&gt;white-hot
potato&lt;/a&gt; that they unwittingly
caught when they first made a good technical choice, this creates a
perverse long-term incentive. &lt;a class=footnote-backref href=#fnref:3:who-is-open-source-about-2026-9 title="Jump back to footnote 3 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:4:who-is-open-source-about-2026-9&gt;
&lt;p id=fn:4&gt;I also find it darkly amusing that there is an explicit affordance here
made for advertising, specifically, “Packages with code that can be used to
display ads are fine. Packages that themselves display ads are not.”  This
distinction rather gives the game away, that this is a website for &lt;a href="https://time.com/archive/6844497/behavior-a-primer-of-american-carnival-talk/"&gt;carnies
and not for
marks&lt;/a&gt;,
and that at some level we expect our users to deserve a lower level of
respect than ourselves.  But a full exploration of that is another blog
post, or maybe a book, that I don’t have time to write right now. &lt;a class=footnote-backref href=#fnref:4:who-is-open-source-about-2026-9 title="Jump back to footnote 4 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:5:who-is-open-source-about-2026-9&gt;
&lt;p id=fn:5&gt;If you want the turbocharged ultra-dramatic version of this problem, make
it open source drivers for an optical prosthesis that lets the users &lt;em&gt;see&lt;/em&gt;
instead of an art app.  That level of immediate physical dependency could
be clarifying.  It does also start to edge into an area where you could say
that biomedical devices ought to be regulated differently, and that’s not
really a “software” problem but a “healthcare” problem and I’d mostly
agree.  Except for the fact that this is a &lt;em&gt;very&lt;/em&gt; short distance away from
&lt;a href="https://tech.lgbt/@xogium/110507457689374019"&gt;breaking everyone’s screen-reader with no notice or
recourse&lt;/a&gt;. &lt;a class=footnote-backref href=#fnref:5:who-is-open-source-about-2026-9 title="Jump back to footnote 5 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:6:who-is-open-source-about-2026-9&gt;
&lt;p id=fn:6&gt;Did you believe I could write a blog post in 2026 which wasn’t somehow
about AI?  I wish I could still believe that. &lt;a class=footnote-backref href=#fnref:6:who-is-open-source-about-2026-9 title="Jump back to footnote 6 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:7:who-is-open-source-about-2026-9&gt;
&lt;p id=fn:7&gt;It doesn’t help that many of the most pro-AI voices are starting to have
an, ahem, &lt;a href="https://www.bbc.com/news/articles/c34gd48x5rlwo"&gt;discernible political
valence&lt;/a&gt; that is very
unpopular among users. &lt;a class=footnote-backref href=#fnref:7:who-is-open-source-about-2026-9 title="Jump back to footnote 7 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:8:who-is-open-source-about-2026-9&gt;
&lt;p id=fn:8&gt;My apologies to MLK. &lt;a class=footnote-backref href=#fnref:8:who-is-open-source-about-2026-9 title="Jump back to footnote 8 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:9:who-is-open-source-about-2026-9&gt;
&lt;p id=fn:9&gt;If you read this whole post you can see that I sure need the help with
all that. &lt;a class=footnote-backref href=#fnref:9:who-is-open-source-about-2026-9 title="Jump back to footnote 9 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;&lt;/body&gt;</content><category term="misc"></category><category term="open-source"></category><category term="software"></category><category term="ai"></category></entry><entry><title>... but what about video games?</title><link href="https://blog.glyph.im/2026/09/but-what-about-video-games.html" rel="alternate"></link><published>2026-09-06T15:57:00-07:00</published><updated>2026-09-06T15:57:00-07:00</updated><author><name>Glyph</name></author><id>tag:blog.glyph.im,2026-09-06:/2026/09/but-what-about-video-games.html</id><summary type="html">&lt;p&gt;Playing video games also uses a GPU. Is using a local LLM for coding any worse than that?&lt;/p&gt;</summary><content type="html">&lt;body&gt;&lt;p&gt;I get asked this rhetorical question a lot, in various forms:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Sure, datacenters might use a lot of energy, but you don’t &lt;em&gt;have&lt;/em&gt; to use a
hosted frontier model to do software development.  What if I just run a local
open-weights model to do some coding, with an open-source coding agent?
Video games also use my GPU.  Is local model development any worse than
playing a video game?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;So I want to write down my comprehensive answer to this: Yes, using an LLM to
write some code is worse than playing a video game, for a few reasons.&lt;/p&gt;
&lt;h2 id=video-games-are-interactive-llms-are-batch-jobs&gt;Video Games Are Interactive, LLMs Are Batch Jobs&lt;/h2&gt;
&lt;p&gt;Video games use compute to respond to human input. You are using your GPU while
you are looking at a screen, displaying an image. When you are done playing,
you shut off the game, and your computer goes back to idle. It’s much less
energy. By contrast, agentic loops with evals (the only kind of “AI” that is
meaningfully any good at coding) are running hot, for days. To use the most
recent example of such a thing, &lt;a href="https://forums.paint.net/topic/134563-🍷-extremely-experimental-winelinux-support-how-to-get-started/"&gt;a &lt;em&gt;very rough&lt;/em&gt; first sketch of an
implementation of a Windows graphics API backend to help port a paint program
to other
platforms&lt;/a&gt;,
it took 3 weeks of Claude time, “day and night”. Do you play a lot of video
games for 500 hours to make it past the tutorial level, while &lt;em&gt;also&lt;/em&gt; using
&lt;em&gt;other&lt;/em&gt; computers for other things, as well as the rest of your carbon
footprint?&lt;/p&gt;
&lt;h2 id=video-games-need-development-llms-need-training&gt;Video Games Need Development, LLMs Need Training&lt;/h2&gt;
&lt;p&gt;Video games use compute to respond to human input during development, too.
Your game has to be made, but your LLM has to be &lt;em&gt;trained&lt;/em&gt;.  LLMs use a
historically extreme amount of power, &lt;em&gt;probably&lt;/em&gt; using &lt;a href="https://science.feedback.org/training-and-using-chatgpt-uses-a-lot-of-energy-but-exact-numbers-are-tricky-to-pin-down-without-data-from-openai/"&gt;more than the entire
Internet&lt;/a&gt;,
but it’s kind of hard to say.  Still, it seems a reasonable estimate to within
several orders of magnitude that even over a multi-year project with hundreds
of developers, the power used to develop an individual video game is &lt;em&gt;nowhere
close&lt;/em&gt; to training even a small LLM.&lt;/p&gt;
&lt;p&gt;This is true even for local models.  OpenAI has &lt;a href="https://www.fdd.org/analysis/2026/02/13/openai-alleges-chinas-deepseek-stole-its-intellectual-property-to-train-its-own-models/"&gt;openly claimed that DeepSeek
“stole its intellectual
property”&lt;/a&gt;,
and I have heard grumblings that none of the open-weights generalist models
could realistically exist without the massive lift that the frontier labs are
doing with their training, in various other ways too.  Secrecy throughout the
industry makes this kind of impossible to understand rigorously, but it seems
fair to say that you are partially culpable for all that famously
energy-intensive frontier lab training if you’re using a local model.&lt;/p&gt;
&lt;h3 id=and-they-keep-needing-training&gt;And They &lt;em&gt;Keep&lt;/em&gt; Needing Training&lt;/h3&gt;
&lt;p&gt;You also can’t dismiss this as a sunk cost, because in order to stay current
with industry developments, models need to be updated with new information from
the rest of the world, which means that you need to &lt;em&gt;keep&lt;/em&gt; training them.
Beyond the energy for your own use, if you want a real-life agentic workflow
that actually does useful stuff, practically speaking you would still need to
update your local models over and over again, at least once every few months,
which means you would be incentivizing continued energy consumption by whoever
was doing that training for you, including the energy cost of scraping.&lt;/p&gt;
&lt;h2 id=lets-be-real-here-you-arent-actually-using-a-local-model&gt;Let’s Be Real Here, You Aren’t Actually Using A Local Model&lt;/h2&gt;
&lt;p&gt;This question is a hypothetical thought experiment.  Despite &lt;a href="https://www.faros.ai/blog/open-models-vs-frontier-models"&gt;synthetic
benchmarks that keep showing there isn’t much
difference&lt;/a&gt; between
open weight and frontier models, &lt;a href="https://aimultiple.com/llm-market-share"&gt;nobody’s actually using local models for much
of anything&lt;/a&gt; beyond sharing those
talking points.  Depending on which benchmark you’re looking at, &lt;a href="https://whatllm.org/blog/open-source-vs-proprietary-llms-2026"&gt;maybe it’s
good enough or maybe it’s
worse&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;As an inveterate AI hater, all these systems seem pretty bad to me, but it
seems that people who find them useful tend to &lt;em&gt;subjectively&lt;/em&gt; believe the
frontier models are worth the premium, and that’s what they’re actually using.
Once you have accepted that it is OK to use LLMs for coding at all, it seems
like a &lt;em&gt;very&lt;/em&gt; quick slippery slope on down to “we’ll go ahead and use the
frontier models for now anyway, but we could be ethically better in the future
by switching to an open weights one, that option is always available”.&lt;/p&gt;
&lt;h2 id=theres-a-reason-we-have-data-centers&gt;There’s A Reason We Have Data Centers&lt;/h2&gt;
&lt;p&gt;Devolving power usage to local LLMs might be good to make users responsible for
their costs and decrease the impacts to communities that are physically next to
huge concentrations of power utilization, not to mention generation.  However,
there’s a reason that it makes sense for the providers to build these giant
facilities: economies of scale &lt;em&gt;reduce&lt;/em&gt; total power consumption, they don’t
increase it.  If you do all the same stuff with a local model that they have to
do in hosted environments, &lt;a href="https://omniforge.online/blog/green-ai-at-the-edge"&gt;it will probably take &lt;em&gt;more&lt;/em&gt; power, even though you
will be incentivized to do different
stuff&lt;/a&gt;.  This incentive to
“do different stuff” is why although local models can hypothetically hold their
own against the frontier labs for some tasks, when people or businesses take
their inference costs in-house they often find that it’s too painful and move
back to hosted LLMs.&lt;/p&gt;
&lt;h2 id=there-are-problems-other-than-power&gt;There Are Problems Other Than Power&lt;/h2&gt;
&lt;p&gt;These are subjects for a different post, but you have to consider a lot of
other externalities: AI psychosis, de-skilling, comprehension debt, cultivating
a dependency, introducing security defects, limiting your design space based on
what LLMs can understand, context rot, wasting time on invalid solutions,
introducing unpredictability into your workflows.  You still have to consider
the &lt;a href="https://blog.glyph.im/2025/08/futzing-fraction.html"&gt;total cost benefit ratio&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=to-sum-up&gt;To Sum Up&lt;/h2&gt;
&lt;p&gt;Local LLMs might alleviate &lt;em&gt;some&lt;/em&gt; of the harms from using the hosted frontier
providers.  There are fewer privacy concerns, you can measure your power
utilization and be more directly responsible for it, you can build interfaces
with affordances that are less oriented towards addiction and dependency than
the major frontier labs’ harnesses.&lt;/p&gt;
&lt;p&gt;But they’re not automatically “the same as playing a video game” just because
they can use the same GPU.&lt;/p&gt;
&lt;h2 id=acknowledgments&gt;Acknowledgments&lt;/h2&gt;
&lt;p class=update-note&gt;Thank you to &lt;a href="/pages/patrons.html"&gt;my patrons&lt;/a&gt; who are supporting my writing on
this blog.  If you like what you’ve read here and you’d
like to read more of it, or you’d like to support my &lt;a href="https://github.com/glyph/"&gt;various open-source
endeavors&lt;/a&gt;, you can &lt;a href="/pages/patrons.html"&gt;support my work as a
sponsor&lt;/a&gt;!&lt;/p&gt;&lt;/body&gt;</content><category term="misc"></category><category term="ai"></category><category term="llm"></category><category term="programming"></category></entry><entry><title>Adversarial Communication</title><link href="https://blog.glyph.im/2026/06/adversarial-communication.html" rel="alternate"></link><published>2026-06-23T13:06:00-07:00</published><updated>2026-06-23T13:06:00-07:00</updated><author><name>Glyph</name></author><id>tag:blog.glyph.im,2026-06-23:/2026/06/adversarial-communication.html</id><summary type="html">&lt;p&gt;“AI” turns every conversation into a fight, because fighting is what
they are good at.&lt;/p&gt;</summary><content type="html">&lt;body&gt;&lt;p&gt;As I have discussed in &lt;a href="https://blog.glyph.im/2025/08/futzing-fraction.html"&gt;previous posts&lt;/a&gt;,
&lt;em&gt;“AIs” can make mistakes&lt;/em&gt;.  In fact, they &lt;em&gt;do&lt;/em&gt; make mistakes, and their
mistake-making patterns are such that where and how they will make mistakes is
both uncertain and constantly changing.&lt;/p&gt;
&lt;p&gt;Thus, in any scenario where you want to attempt to make “productive” use of
“AI”, you must have a system in place for checking every result.  Not checking
&lt;em&gt;some&lt;/em&gt; results; checking &lt;em&gt;every&lt;/em&gt; result.  If each result might have a
consequence for you (and if it didn’t have a consequence, why bother automating
it?) and you cannot predict in advance which kinds of results will need
verification, then verification is always required.&lt;/p&gt;
&lt;p&gt;The verification often ends up being just as expensive as doing the work in the
first place, which means that if you want your usage of “AI” to be personally
profitable, you have to find someone &lt;em&gt;else&lt;/em&gt; to externalize the cost of
verification onto.  This person becomes your adversary, and, if you are
successful, your “AI’s” victim.&lt;/p&gt;
&lt;h2 id=the-ladder-climber-and-their-reverse-centaur-rungs&gt;The Ladder-Climber And Their Reverse-Centaur Rungs&lt;/h2&gt;
&lt;p&gt;One way that this constellation of facts can straightforwardly assemble
themselves into a dystopian nightmare is the phenomenon, described by Cory
Doctorow, of the &lt;a href="https://locusmag.com/feature/commentary-cory-doctorow-reverse-centaurs/"&gt;reverse
centaur&lt;/a&gt;.
This is when your employer non-consensually turns &lt;em&gt;you&lt;/em&gt; into the verification
system.  The “AI” does the fun part of initially performing the work, and then
you do the boring part where you check if the robot is right and clean up its
messes, even if &lt;a href="https://fortune.com/article/why-is-the-cost-of-ai-higher-than-human-workers-nvidia-executive/"&gt;everyone already knows that it would, in aggregate, be cheaper
for you to do the work in the first
place&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Reverse centaurs can be made from any automation, not &lt;em&gt;only&lt;/em&gt; “AI” automation.
I think that there is a reason that this term happens to have emerged in the
“age of AI”, though, and not with earlier automation technologies (even those
which were
&lt;a href="https://en.wikipedia.org/wiki/Cotton_gin#Effects_in_the_United_States"&gt;considerably&lt;/a&gt;
more &lt;a href="https://victorianweb.org/technology/ir/capuano.html"&gt;viscerally
horrific&lt;/a&gt;).  That reason
is: the &lt;em&gt;wrongness&lt;/em&gt; of “AI” output is not merely a technical feature that must
be compensated for, it is a generalized externality.&lt;/p&gt;
&lt;p&gt;As I mentioned above, if you are responsible for the entirety of the work, both
extruding the “AI” output &lt;em&gt;and&lt;/em&gt; checking it, it’s usually cheaper to have
humans do the entirety of the work to begin with.  When humans do the writing
directly, we can check as we go, and thus verification doesn’t need to be as
comprehensive.&lt;/p&gt;
&lt;p&gt;When “AI” coding advocates say “code review is the
&lt;a href="https://lucumr.pocoo.org/2026/2/13/the-final-bottleneck/"&gt;bottleneck&lt;/a&gt;”, what
they are observing is that the LLM is still rolling the dice for each PR, and a
human is still necessary to verify that each of those rolls is a winner.  But
calling this process “code review” is a bit of a
&lt;a href="https://blog.glyph.im/2026/03/what-is-code-review-for.html"&gt;misnomer&lt;/a&gt;; it’s not really “code
review” in the traditional sense, it’s &lt;em&gt;human understanding&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;Before the advent of “AI”, the human understanding was implicit in the process
of writing the code in the first place&lt;sup id=fnref:1:adversarial-communication-2026-6&gt;&lt;a class=footnote-ref href=#fn:1:adversarial-communication-2026-6 id=fnref:1&gt;1&lt;/a&gt;&lt;/sup&gt;, and the code review was a way of
diffusing and extending that understanding.  Now that the code can be authored
with no initial understanding taking place, that cost has not gone away, it has
moved.&lt;/p&gt;
&lt;p&gt;Human understanding was &lt;em&gt;always&lt;/em&gt; the bottleneck.&lt;/p&gt;
&lt;p&gt;However, this is taking a &lt;em&gt;collaborative&lt;/em&gt; view of a software project, where
satisfying the needs and solving the problems of your customers are the goals.
We can see that “AI” is a bad tool to satisfy those goals, because all it’s
doing is converting the first half of the work, that of understanding the code
as you write it, to understanding the agent’s output as you read it.&lt;/p&gt;
&lt;p&gt;What if, instead, we were to take the view that every software company is a
Hobbesian nightmare, red in tooth and claw? In this view, the only goal of a
software project is for the individual developers to make their promo cycles
and get their bonuses.  Given that there is only a certain amount of money to
go around, this is a zero-sum game where each programmer wants to look more
productive than their colleagues.&lt;/p&gt;
&lt;p&gt;Pretty much &lt;em&gt;every&lt;/em&gt; organization finds it easy to reward “productivity” as
expressed by lines of code emitted, but the benefits of doing
thorough and thoughtful design, analysis, and code review &lt;em&gt;very difficult&lt;/em&gt; to
reward.  In this world, an LLM is an invaluable tool for the sociopathic
ladder-climber, particularly if your legacy organization is still structuring
their workflows as if the person prompting the bot is “writing” the code, and
then they get to foist off the act of “reviewing” the code onto someone else.&lt;/p&gt;
&lt;p&gt;Here, the prompter effectively externalizes the cost of the LLM’s failures but
internalizes any benefits.  The prompter will vibe-code a big feature, so large
that the assigned reviewer can’t possibly comprehend it all effectively. When
this happens, the reviewer will, &lt;em&gt;eventually&lt;/em&gt;, be pressured to approve it, even
if they can try to spot a few problems along the way.  The reviewer has their
own work to get back to, after all, the obligation to review the prompter’s
(read: the bot’s) code is a drain on their time that they are not going to get
rewarded for.&lt;/p&gt;
&lt;p&gt;If this feature is a big success, the prompter gets a promotion.  If it causes
a big issue, well, the reviewer must not have been careful enough.&lt;/p&gt;
&lt;p&gt;This is why LLMs are “good for coding”, and also why their biggest promoters
&lt;a href="https://ap7i.com/posts/github-outages-vibe-coding-era/"&gt;keep&lt;/a&gt;
&lt;a href="https://www.linkedin.com/posts/paulsf_last-weeks-massive-google-cloud-outage-activity-7340015235321278466-c0hz"&gt;having&lt;/a&gt;
&lt;a href="https://www.ft.com/content/7cab4ec7-4712-4137-b602-119a44f771de"&gt;outages&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=the-generative-gish-galloper&gt;The Generative Gish Galloper&lt;/h2&gt;
&lt;p&gt;Coding is the biggest “success story” of this type of adversarial
communication, but it is by far not the only instance of such a thing.  LLMs
create a new form of leverage that can turn &lt;a href="https://en.wikipedia.org/wiki/Brandolini%27s_law"&gt;Brandolini’s
law&lt;/a&gt; from a linear advantage
into an exponential one.  If you are engaged in a political debate where you
want to overwhelm the other side in nonsense, an LLM can generate bullshit
faster than it is physically possible for a human being to type, let alone
respond thoughtfully.  There is an asymmetry to the utility of this weapon as
well: only one side of the political spectrum wants to &lt;a href="https://en.wikipedia.org/wiki/Flood_the_zone"&gt;flood the
zone&lt;/a&gt; and destroy trust in
institutions and the concept of truth.  There’s a good reason that &lt;a href="https://newsocialist.org.uk/transmissions/ai-the-new-aesthetics-of-fascism/"&gt;the
fascists love
it&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=straightforward-spam-and-fraud&gt;Straightforward Spam and Fraud&lt;/h2&gt;
&lt;p&gt;This is kind of obvious, but LLMs can generate lightly-customized,
plausible-looking text much more quickly than any human being.  &lt;a href="https://withpersona.com/blog/llm-fraud"&gt;This
facilitates their use in fraud, spam, and
scams.&lt;/a&gt; In a spamming or fraudulent
interaction, once again, the costs are externalized onto the victim: the
recipient of a spam message has to do all the work of “checking” the LLM’s
output.  Spammers already expect very low hit rates from boilerplate, and if the
LLM can increase those percentages from 1% to 5% the technology will pay for
itself; they don’t need anything like &lt;em&gt;reliable&lt;/em&gt; accuracy.&lt;/p&gt;
&lt;h2 id=customer-support&gt;Customer “Support”&lt;/h2&gt;
&lt;p&gt;If you have any kind of commercial relationship with a company, I probably
don’t even need to mention this: customer “support” bots are a misery.
&lt;a href="https://www.forbes.com/sites/terdawn-deboe/2026/04/20/customers-hate-your-ai-chatbot-small-businesses-should-listen/"&gt;Everybody knows
it&lt;/a&gt;
at this point.  But customer support is usually conceptualized by businesses as
an adversarial interaction, because it is a cost center.  They maintain
internal metrics on time-to-resolution and try to optimize them.  Implicitly,
this creates a dynamic where the goal of the customer service agent’s job is
not to solve your problem, but to emit noise that will cause you to &lt;em&gt;think&lt;/em&gt;
your problem is resolved, or to give up, as fast as possible.  Unsurprisingly,
LLMs can emit this noise faster than humans can, getting those customers off
the phone.  But those customers will &lt;em&gt;remember&lt;/em&gt; those interactions, and the
story &lt;em&gt;outside&lt;/em&gt; the TTR metrics is horrible.&lt;/p&gt;
&lt;p&gt;Similarly to the situation in software development, LLMs can look very good on
paper for customer support, but mostly what they are doing is illuminating the
problems with the industry’s existing metrics, by turning “winning the metrics
battle against the customer” into a more obvious and immediate defeat for the
company’s long term reputation.&lt;/p&gt;
&lt;h2 id=education&gt;“Education”&lt;/h2&gt;
&lt;p&gt;In 2026 it is sadly a fact of life that &lt;a href="https://www.nytimes.com/2026/06/18/us/ai-apps-students-cheat.html"&gt;students cheat all the time using
“AI”&lt;/a&gt;, and
that this cheating is very successful, in that the teachers find it very hard
to detect.&lt;/p&gt;
&lt;p&gt;LLMs are great for cheating on schoolwork because the student is externalizing
the work of the checking onto the teachers, who are often starting at a
disadvantage to begin with, at least in the US.&lt;/p&gt;
&lt;p&gt;My view is that this is happening because of a divergence in the way that
students vs. teachers (or, more accurately, “the broader educational system”)
view grading.&lt;/p&gt;
&lt;p&gt;When a student is asked to write an essay, the teachers see the effort as both
intrinsically worthwhile for the student, as well as useful as a pedagogical
tool to evaluate and react to the student’s progress.  The student, by
contrast, sees a stumbling block designed to knock them off the path to success
and into a permanent underclass.  It is no wonder that the student sees “AI” as
useful to their own goals and has no compunction about deploying it.&lt;/p&gt;
&lt;p&gt;There is a bitter irony that the ability to understand the inherent value of
actually writing the essay on their own is the sort of thing that students can
really only learn by writing a bunch of essays.  There’s no way that I can
think of which makes the benefit legible as long as a shortcut is available.&lt;/p&gt;
&lt;p&gt;The net effect here is a downward spiral, where the already-wobbling
educational system is sustaining an attack that it doesn’t have the resources
to recover from.  The individual students’ attacks against their teachers and
their schools’ grading systems might appear to momentarily succeed, but they
will win the battle and lose the war.&lt;/p&gt;
&lt;h2 id=spamming-for-good&gt;Spamming “For Good”?&lt;/h2&gt;
&lt;p&gt;Usually when we talk about someone unilaterally choosing to enter into an
adversarial relationship, that’s an “attack” and for good reasons we have a
negative impression of the attacker.  However, I would be remiss if I did not
point out that there are some cases where the relationship was already
adversarial; just because you’re the attacker doesn’t mean that you are evil.&lt;/p&gt;
&lt;p&gt;For example we might &lt;em&gt;imagine&lt;/em&gt; use-cases like automatically filing appeals for
prior authorizations against health insurance. It’s relatively
&lt;a href="https://en.wikipedia.org/wiki/Delay,_Deny,_Defend"&gt;well-known&lt;/a&gt; at this point
that the main way for-profit insurers maintain their margins is by denying
claims right up to the line of the policies themselves being fraud, so using a
spamming tool to fight them might be entirely justifiable&lt;sup id=fnref:2:adversarial-communication-2026-6&gt;&lt;a class=footnote-ref href=#fn:2:adversarial-communication-2026-6 id=fnref:2&gt;2&lt;/a&gt;&lt;/sup&gt; in that case.&lt;/p&gt;
&lt;p&gt;Similarly, using an LLM could be justified in a fight against a company
refusing to honor a warranty.  One could imagine using an LLM to immediately
generate replies and escalations.&lt;/p&gt;
&lt;p&gt;However, even in imagined cases like these, the underlying problem is that the
insurers and the vendors already have a tremendous amount of structural power,
so it is more likely that they will have the advantage in deploying a
communications weapon like an LLM, as well as enacting policies to simply
ignore any LLM-based communication that you might submit.  Worse, if these
strategies were to become widespread, they might provide an excuse to reject
&lt;em&gt;any&lt;/em&gt; communications by feeding them into an unreliable “&lt;a href="https://www.npr.org/2025/12/16/nx-s1-5492397/ai-schools-teachers-students"&gt;LLM
detector&lt;/a&gt;”
and issuing an automated “computer says no” even to hand-written
correspondence.&lt;/p&gt;
&lt;p&gt;It is also worth stressing that these cases are imagined, as compared to the
very real coworker-abuse, spam, scam, fraud, and disinformation campaigns being
waged in real life today.&lt;/p&gt;
&lt;p&gt;Therefore, while legitimate uses might exist, it’s hard to imagine that there’s
anywhere they would be genuinely valuable and sustainable.  In the best case
“AI” will provide a temporary advantage for underdogs that will provoke an arms
race which the resource-advantaged adversaries will win in the long run, in the
worst case the arms race itself will cement permanent structural change that
will make things worse.&lt;/p&gt;
&lt;h2 id=search-by-stealing&gt;“Search” By Stealing&lt;/h2&gt;
&lt;p&gt;Most of the adversarial utility of “AI” is on the “write” side, since
write-amplification is more obviously aggressive than reading.  But the “read”
side of LLMs — summarization and question-answering — can be a form of attack
as well.&lt;/p&gt;
&lt;p&gt;To begin with, &lt;a href="https://www.theregister.com/software/2025/08/29/ai-crawlers-destroying-websites-in-hunger-for-content/464120"&gt;the act of reading
itself&lt;/a&gt;
is currently enormously destructive, but that’s arguably not a &lt;em&gt;fundamental&lt;/em&gt;
aspect of this technology.  They &lt;em&gt;could&lt;/em&gt; set reasonable rate-limits and respect
things like &lt;code&gt;robots.txt&lt;/code&gt;, as search engines have for decades now.  They could
also refrain from committing &lt;a href="https://www.npr.org/2025/09/05/nx-s1-5529404/anthropic-settlement-authors-copyright-ai"&gt;criminal
levels&lt;/a&gt;
&lt;a href="https://www.theguardian.com/technology/2025/jan/10/mark-zuckerberg-meta-books-ai-models-sarah-silverman"&gt;of copyright
infringement&lt;/a&gt;.
But, today, using “AI” tools does suborn this sort of out-of-control crawling.&lt;/p&gt;
&lt;p&gt;More insidiously, consider the scenario described in &lt;a href="https://www.youtube.com/watch?v=8KQFgWdiudo"&gt;this YouTube
video&lt;/a&gt;.  The LTT Bros decided to
try Linux again, and in the course of so doing, they had problems.  When trying
to solve these problems, they were faced with a choice: they could consult
Reddit, or they could ask an LLM.  Asking an LLM would “gaslight the heck out
of” them, but they still found it preferable, because they would at least get
an answer without getting yelled at.&lt;/p&gt;
&lt;p&gt;Initially this sounds great.  But it also means that you want to extract
knowledge from a community, while mechanically eliding any values or norms that
the community may want to impart as part of offering that knowledge.  As
someone who spent many years in a community tech support role, this is
worrying.  Many requests for support are people asking how to do things that
will momentarily solve a superficial problem but create a long-term reliability
problem or even an immediate security risk, that the question-asker doesn’t
want to hear about.  Consider the question “I’m tired of entering my password
so much, how do I make it so my laptop unlocks automatically”.  An obsequious
chatbot will helpfully tell you how to do this without pushback.&lt;/p&gt;
&lt;p&gt;But, this is also a sort of ethically murky area.  The Linux community is
somewhat famously, for &lt;a href="https://news.ycombinator.com/item?id=10332286"&gt;many years
now&lt;/a&gt;, a toxic cesspool of
general hostility, misogyny, etc.  It is certainly a good thing that people can
get access to this knowledge without subjecting themselves to abuse.  But it
also means that the people &lt;em&gt;with&lt;/em&gt; the power and the privilege to change the
community for the better can just quietly withdraw, rather than fixing the
problems.  It also means that the positive elements of culture cannot be
transmitted, and people will have no opportunity to learn about unknown
unknowns.&lt;/p&gt;
&lt;p&gt;In this case, the “adversarial” communication is with society.  The thing that
using an LLM for search lets you do is withdraw from society and avoid forming
any personal connections.  There are some personal connections which are
painful and annoying, and so that can feel like a momentary balm.  But the need
to make connections &lt;em&gt;in general&lt;/em&gt; is, like, the concept of society itself.&lt;/p&gt;
&lt;h2 id=who-am-i-hurting&gt;Who Am I Hurting?&lt;/h2&gt;
&lt;p&gt;LLMs are good at adversarial communication.  They are &lt;em&gt;so&lt;/em&gt; good at it, relative
to their other benefits, that they will tend to &lt;em&gt;make&lt;/em&gt; communications
adversarial if you are not remaining vigilant about the possibility that it
might do so.  My request to you, dear reader, if you are going to use such
tools, is to always ask yourself, “who might I be hurting, if I use an LLM for
this?”&lt;/p&gt;
&lt;p&gt;If you’re using an “AI”, who is its adversary?  If you haven’t given it one
yet, who might the “AI” &lt;em&gt;turn into&lt;/em&gt; an adversary?  Who might you overwhelm with
an asymmetric amount of output, or, if you’re receiving information and not
sending it, who are you taking that information from without consulting?&lt;/p&gt;
&lt;p&gt;Figure out the answers to these questions and conduct yourself accordingly; the
answer might be “yourself”.&lt;/p&gt;
&lt;h2 id=acknowledgments&gt;Acknowledgments&lt;/h2&gt;
&lt;p class=update-note&gt;Thank you to &lt;a href="/pages/patrons.html"&gt;my patrons&lt;/a&gt; who are supporting my writing on
this blog.  If you like what you’ve read here and you’d
like to read more of it, or you’d like to support my &lt;a href="https://github.com/glyph/"&gt;various open-source
endeavors&lt;/a&gt;, you can &lt;a href="/pages/patrons.html"&gt;support my work as a
sponsor&lt;/a&gt;!&lt;/p&gt;
&lt;div class=footnote&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=fn:1:adversarial-communication-2026-6&gt;
&lt;p id=fn:1&gt;One of the reasons that software developers tend to prefer
&lt;a href="https://en.wikipedia.org/wiki/Greenfield_project"&gt;greenfield&lt;/a&gt; development
is that when you are given a blank page, you can project your &lt;em&gt;own&lt;/em&gt;
specific understanding onto it.  You can structure the codebase in a way
that works for your brain, down to the variable naming conventions and the
module layouts.  LLM-assisted development makes everything into instant
brownfield work, which makes developers instantly miserable; even those who
are excited about the technology will frequently complain about how it
feels like their agency has been stolen and their joy in the work has been
diminished.  But I digress. &lt;a class=footnote-backref href=#fnref:1:adversarial-communication-2026-6 title="Jump back to footnote 1 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:2:adversarial-communication-2026-6&gt;
&lt;p id=fn:2&gt;Modulo the massive amount of &lt;em&gt;other&lt;/em&gt; externalities involved in using
LLMs, of course, but I don’t have the time or energy to get into those
here. &lt;a class=footnote-backref href=#fnref:2:adversarial-communication-2026-6 title="Jump back to footnote 2 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;&lt;/body&gt;</content><category term="misc"></category><category term="ai"></category><category term="llm"></category><category term="programming"></category></entry><entry><title>Opaque Types in Python</title><link href="https://blog.glyph.im/2026/05/opaque-types-in-python.html" rel="alternate"></link><published>2026-05-21T17:33:00-07:00</published><updated>2026-05-21T17:33:00-07:00</updated><author><name>Glyph</name></author><id>tag:blog.glyph.im,2026-05-21:/2026/05/opaque-types-in-python.html</id><summary type="html">&lt;p&gt;A proposed technique for exposing an opaque data structure with
idiomatic modern Python.&lt;/p&gt;</summary><content type="html">&lt;body&gt;&lt;p&gt;Let’s say you’re writing a Python library.&lt;/p&gt;
&lt;p&gt;In this library, you have some collection of state that represents “options” or
“configuration” for a bunch of operations.  Such a set of options is a bundle
of potentially ever-increasing complexity.  Thus, you will want it to have an
extremely minimal compatibility surface, with a very carefully chosen public
interface, that is either small, or perhaps nothing at all.  Such an object
conveys state and might have some private behavior, but all you want consumers
to be able to do is build it in very constrained, specific ways, and then pass
it along as a parameter to your own APIs.&lt;/p&gt;
&lt;p&gt;By way of example, imagine that you’re wrapping a library that handles shipping
physical packages.&lt;/p&gt;
&lt;p&gt;There are a zillion ways to do it ship a package.  There are different carriers
who can ship it for you. There’s air freight, and ground freight, and sea
freight.  There’s overnight shipping.  There’s the option to require a
signature.  There’s package tracking and certified mail.  Suffice it to say,
lots of stuff.&lt;/p&gt;
&lt;p&gt;If you are starting out to implement such a library, you might need an object
called something like &lt;code&gt;ShippingOptions&lt;/code&gt; that encapsulates some of this.  At the
core of your library you might have a function like this:&lt;/p&gt;
&lt;div class=highlight&gt;&lt;table class=highlighttable&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class=linenos&gt;&lt;div class=linenodiv&gt;&lt;pre&gt;&lt;span class=normal&gt;1&lt;/span&gt;
&lt;span class=normal&gt;2&lt;/span&gt;
&lt;span class=normal&gt;3&lt;/span&gt;
&lt;span class=normal&gt;4&lt;/span&gt;
&lt;span class=normal&gt;5&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=code&gt;&lt;div&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class=k&gt;async&lt;/span&gt; &lt;span class=k&gt;def&lt;/span&gt; &lt;span class=nf&gt;shipPackage&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;
        &lt;span class=n&gt;how&lt;/span&gt;&lt;span class=p&gt;:&lt;/span&gt; &lt;span class=n&gt;ShippingOptions&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt;
        &lt;span class=n&gt;where&lt;/span&gt;&lt;span class=p&gt;:&lt;/span&gt; &lt;span class=n&gt;Address&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt;
    &lt;span class=p&gt;)&lt;/span&gt; &lt;span class=o&gt;-&amp;gt;&lt;/span&gt; &lt;span class=n&gt;ShippingStatus&lt;/span&gt;&lt;span class=p&gt;:&lt;/span&gt;
    &lt;span class=o&gt;...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If you are &lt;em&gt;starting out&lt;/em&gt; implementing such a library, you know that you’re
going to get the initial implementation of &lt;code&gt;ShippingOptions&lt;/code&gt; wrong; or, at the
very least, if not “wrong”, then “incomplete”.  You should not want to commit
to an expansive public API with a ton of different attributes until you really
understand the problem domain pretty well.&lt;/p&gt;
&lt;p&gt;Yet, &lt;code&gt;ShippingOptions&lt;/code&gt; is absolutely vital to the rest of your library.  You’ll
need to construct it and pass it to various methods like &lt;code&gt;estimateShippingCost&lt;/code&gt;
and &lt;code&gt;shipPackage&lt;/code&gt;.  So you’re not going to want a ton of complexity and churn
as you evolve it to be more complex.&lt;/p&gt;
&lt;p&gt;Worse yet, this object has to hold a ton of state.  It’s got attributes, maybe
even quite complex internal attributes that relate to different shipping
services.&lt;/p&gt;
&lt;p&gt;Right now, &lt;em&gt;today&lt;/em&gt;, you need to add &lt;em&gt;something&lt;/em&gt; so you can have “no rush”,
“standard” and “expedited” options.  You can’t just put off implementing that
indefinitely until you can come up with the perfect shape. What to do?&lt;/p&gt;
&lt;p&gt;The tool you want here is the &lt;em&gt;opaque data type&lt;/em&gt; design pattern.  C is lousy
with such things (&lt;code&gt;FILE&lt;/code&gt;, &lt;code&gt;pthread_*_t&lt;/code&gt;, &lt;code&gt;fd_set&lt;/code&gt;, etc).  A &lt;code&gt;typedef&lt;/code&gt; in a
header file can easily achieve this.&lt;/p&gt;
&lt;p&gt;But in Python, if you expose a &lt;code&gt;dataclass&lt;/code&gt; — or &lt;em&gt;any&lt;/em&gt; class, really — even if
you keep all your fields private, the &lt;em&gt;constructor&lt;/em&gt; is still, inherently,
public.  You can make it raise an exception or something, but your type checker
still won’t help your users; it’ll still look like it’s a normal class.&lt;/p&gt;
&lt;p&gt;Luckily, Python typing provides a tool for this:
&lt;a href="https://docs.python.org/3.14/library/typing.html#newtype"&gt;&lt;code&gt;typing.NewType&lt;/code&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Let’s review our requirements:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;We need a type that our client code can use in its type annotations; it
   needs to be public.&lt;/li&gt;
&lt;li&gt;They need to be able to consruct it &lt;em&gt;somehow&lt;/em&gt;, even if they shouldn’t be
   able to see its attributes or its internal constructor arguments.&lt;/li&gt;
&lt;li&gt;To express high-level things (like “ship fast”) that should stay supported
   as we add more nuanced and complex configurations in the future (like “ship
   with the fastest possible option provided by the lowest-cost carrier that
   supports signature verification”).&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;In order to solve these problems respectively, we will use:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;a &lt;em&gt;public &lt;code&gt;NewType&lt;/code&gt;&lt;/em&gt;, which gives us our public name...&lt;/li&gt;
&lt;li&gt;which wraps a &lt;em&gt;private class&lt;/em&gt; with entirely private attributes, to give us
   an actual data structure, while not exposing the constructor,&lt;/li&gt;
&lt;li&gt;a set of public constructor functions, which returns our &lt;code&gt;NewType&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;When we put that all together, it looks like this:&lt;/p&gt;
&lt;div class=highlight&gt;&lt;table class=highlighttable&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class=linenos&gt;&lt;div class=linenodiv&gt;&lt;pre&gt;&lt;span class=normal&gt; 1&lt;/span&gt;
&lt;span class=normal&gt; 2&lt;/span&gt;
&lt;span class=normal&gt; 3&lt;/span&gt;
&lt;span class=normal&gt; 4&lt;/span&gt;
&lt;span class=normal&gt; 5&lt;/span&gt;
&lt;span class=normal&gt; 6&lt;/span&gt;
&lt;span class=normal&gt; 7&lt;/span&gt;
&lt;span class=normal&gt; 8&lt;/span&gt;
&lt;span class=normal&gt; 9&lt;/span&gt;
&lt;span class=normal&gt;10&lt;/span&gt;
&lt;span class=normal&gt;11&lt;/span&gt;
&lt;span class=normal&gt;12&lt;/span&gt;
&lt;span class=normal&gt;13&lt;/span&gt;
&lt;span class=normal&gt;14&lt;/span&gt;
&lt;span class=normal&gt;15&lt;/span&gt;
&lt;span class=normal&gt;16&lt;/span&gt;
&lt;span class=normal&gt;17&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=code&gt;&lt;div&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class=kn&gt;from&lt;/span&gt; &lt;span class=nn&gt;dataclasses&lt;/span&gt; &lt;span class=kn&gt;import&lt;/span&gt; &lt;span class=n&gt;dataclass&lt;/span&gt;
&lt;span class=kn&gt;from&lt;/span&gt; &lt;span class=nn&gt;typing&lt;/span&gt; &lt;span class=kn&gt;import&lt;/span&gt; &lt;span class=n&gt;Literal&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt; &lt;span class=n&gt;NewType&lt;/span&gt;

&lt;span class=nd&gt;@dataclass&lt;/span&gt;
&lt;span class=k&gt;class&lt;/span&gt; &lt;span class=nc&gt;_RealShipOpts&lt;/span&gt;&lt;span class=p&gt;:&lt;/span&gt;
    &lt;span class=n&gt;_speed&lt;/span&gt;&lt;span class=p&gt;:&lt;/span&gt; &lt;span class=n&gt;Literal&lt;/span&gt;&lt;span class=p&gt;[&lt;/span&gt;&lt;span class=s2&gt;"fast"&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt; &lt;span class=s2&gt;"normal"&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt; &lt;span class=s2&gt;"slow"&lt;/span&gt;&lt;span class=p&gt;]&lt;/span&gt;

&lt;span class=n&gt;ShippingOptions&lt;/span&gt; &lt;span class=o&gt;=&lt;/span&gt; &lt;span class=n&gt;NewType&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=s2&gt;"ShippingOptions"&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt; &lt;span class=n&gt;_RealShipOpts&lt;/span&gt;&lt;span class=p&gt;)&lt;/span&gt;

&lt;span class=k&gt;def&lt;/span&gt; &lt;span class=nf&gt;shipFast&lt;/span&gt;&lt;span class=p&gt;()&lt;/span&gt; &lt;span class=o&gt;-&amp;gt;&lt;/span&gt; &lt;span class=n&gt;ShippingOptions&lt;/span&gt;&lt;span class=p&gt;:&lt;/span&gt;
    &lt;span class=k&gt;return&lt;/span&gt; &lt;span class=n&gt;ShippingOptions&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=n&gt;_RealShipOpts&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=s2&gt;"fast"&lt;/span&gt;&lt;span class=p&gt;))&lt;/span&gt;

&lt;span class=k&gt;def&lt;/span&gt; &lt;span class=nf&gt;shipNormal&lt;/span&gt;&lt;span class=p&gt;()&lt;/span&gt; &lt;span class=o&gt;-&amp;gt;&lt;/span&gt; &lt;span class=n&gt;ShippingOptions&lt;/span&gt;&lt;span class=p&gt;:&lt;/span&gt;
    &lt;span class=k&gt;return&lt;/span&gt; &lt;span class=n&gt;ShippingOptions&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=n&gt;_RealShipOpts&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=s2&gt;"normal"&lt;/span&gt;&lt;span class=p&gt;))&lt;/span&gt;

&lt;span class=k&gt;def&lt;/span&gt; &lt;span class=nf&gt;shipSlow&lt;/span&gt;&lt;span class=p&gt;()&lt;/span&gt; &lt;span class=o&gt;-&amp;gt;&lt;/span&gt; &lt;span class=n&gt;ShippingOptions&lt;/span&gt;&lt;span class=p&gt;:&lt;/span&gt;
    &lt;span class=k&gt;return&lt;/span&gt; &lt;span class=n&gt;ShippingOptions&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=n&gt;_RealShipOpts&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=s2&gt;"slow"&lt;/span&gt;&lt;span class=p&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;As a snapshot in time, this is not all that interesting; we could have just
exposed &lt;code&gt;_RealShipOpts&lt;/code&gt; as a public class and saved ourselves some time.  The
fact that this exposes a constructor that takes a string is not a big deal for
the present moment.  For an initial quick and dirty implementation, we can just
do checks like &lt;code&gt;if options._speed == "fast"&lt;/code&gt; in our shipping and estimation
code.&lt;/p&gt;
&lt;p&gt;However, the main thing we are doing here is preserving our flexibility to
evolve the related APIs into the future, so let’s see how we might do that.
For example, let’s allow the shipping options to contain a concrete and
specific carrier and freight method:&lt;/p&gt;
&lt;div class=highlight&gt;&lt;table class=highlighttable&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class=linenos&gt;&lt;div class=linenodiv&gt;&lt;pre&gt;&lt;span class=normal&gt; 1&lt;/span&gt;
&lt;span class=normal&gt; 2&lt;/span&gt;
&lt;span class=normal&gt; 3&lt;/span&gt;
&lt;span class=normal&gt; 4&lt;/span&gt;
&lt;span class=normal&gt; 5&lt;/span&gt;
&lt;span class=normal&gt; 6&lt;/span&gt;
&lt;span class=normal&gt; 7&lt;/span&gt;
&lt;span class=normal&gt; 8&lt;/span&gt;
&lt;span class=normal&gt; 9&lt;/span&gt;
&lt;span class=normal&gt;10&lt;/span&gt;
&lt;span class=normal&gt;11&lt;/span&gt;
&lt;span class=normal&gt;12&lt;/span&gt;
&lt;span class=normal&gt;13&lt;/span&gt;
&lt;span class=normal&gt;14&lt;/span&gt;
&lt;span class=normal&gt;15&lt;/span&gt;
&lt;span class=normal&gt;16&lt;/span&gt;
&lt;span class=normal&gt;17&lt;/span&gt;
&lt;span class=normal&gt;18&lt;/span&gt;
&lt;span class=normal&gt;19&lt;/span&gt;
&lt;span class=normal&gt;20&lt;/span&gt;
&lt;span class=normal&gt;21&lt;/span&gt;
&lt;span class=normal&gt;22&lt;/span&gt;
&lt;span class=normal&gt;23&lt;/span&gt;
&lt;span class=normal&gt;24&lt;/span&gt;
&lt;span class=normal&gt;25&lt;/span&gt;
&lt;span class=normal&gt;26&lt;/span&gt;
&lt;span class=normal&gt;27&lt;/span&gt;
&lt;span class=normal&gt;28&lt;/span&gt;
&lt;span class=normal&gt;29&lt;/span&gt;
&lt;span class=normal&gt;30&lt;/span&gt;
&lt;span class=normal&gt;31&lt;/span&gt;
&lt;span class=normal&gt;32&lt;/span&gt;
&lt;span class=normal&gt;33&lt;/span&gt;
&lt;span class=normal&gt;34&lt;/span&gt;
&lt;span class=normal&gt;35&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=code&gt;&lt;div&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class=kn&gt;from&lt;/span&gt; &lt;span class=nn&gt;dataclasses&lt;/span&gt; &lt;span class=kn&gt;import&lt;/span&gt; &lt;span class=n&gt;dataclass&lt;/span&gt;
&lt;span class=kn&gt;from&lt;/span&gt; &lt;span class=nn&gt;enum&lt;/span&gt; &lt;span class=kn&gt;import&lt;/span&gt; &lt;span class=n&gt;Enum&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt; &lt;span class=n&gt;auto&lt;/span&gt;
&lt;span class=kn&gt;from&lt;/span&gt; &lt;span class=nn&gt;typing&lt;/span&gt; &lt;span class=kn&gt;import&lt;/span&gt; &lt;span class=n&gt;NewType&lt;/span&gt;

&lt;span class=k&gt;class&lt;/span&gt; &lt;span class=nc&gt;Carrier&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=n&gt;Enum&lt;/span&gt;&lt;span class=p&gt;):&lt;/span&gt;
    &lt;span class=n&gt;FedEx&lt;/span&gt; &lt;span class=o&gt;=&lt;/span&gt; &lt;span class=n&gt;auto&lt;/span&gt;&lt;span class=p&gt;()&lt;/span&gt;
    &lt;span class=n&gt;USPS&lt;/span&gt; &lt;span class=o&gt;=&lt;/span&gt; &lt;span class=n&gt;auto&lt;/span&gt;&lt;span class=p&gt;()&lt;/span&gt;
    &lt;span class=n&gt;DHL&lt;/span&gt; &lt;span class=o&gt;=&lt;/span&gt; &lt;span class=n&gt;auto&lt;/span&gt;&lt;span class=p&gt;()&lt;/span&gt;
    &lt;span class=n&gt;UPS&lt;/span&gt; &lt;span class=o&gt;=&lt;/span&gt; &lt;span class=n&gt;auto&lt;/span&gt;&lt;span class=p&gt;()&lt;/span&gt;

&lt;span class=k&gt;class&lt;/span&gt; &lt;span class=nc&gt;Conveyance&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=n&gt;Enum&lt;/span&gt;&lt;span class=p&gt;):&lt;/span&gt;
    &lt;span class=n&gt;air&lt;/span&gt; &lt;span class=o&gt;=&lt;/span&gt; &lt;span class=n&gt;auto&lt;/span&gt;&lt;span class=p&gt;()&lt;/span&gt;
    &lt;span class=n&gt;truck&lt;/span&gt; &lt;span class=o&gt;=&lt;/span&gt; &lt;span class=n&gt;auto&lt;/span&gt;&lt;span class=p&gt;()&lt;/span&gt;
    &lt;span class=n&gt;train&lt;/span&gt; &lt;span class=o&gt;=&lt;/span&gt; &lt;span class=n&gt;auto&lt;/span&gt;&lt;span class=p&gt;()&lt;/span&gt;

&lt;span class=nd&gt;@dataclass&lt;/span&gt;
&lt;span class=k&gt;class&lt;/span&gt; &lt;span class=nc&gt;_RealShipOpts&lt;/span&gt;&lt;span class=p&gt;:&lt;/span&gt;
    &lt;span class=n&gt;_carrier&lt;/span&gt;&lt;span class=p&gt;:&lt;/span&gt; &lt;span class=n&gt;Carrier&lt;/span&gt;
    &lt;span class=n&gt;_freight&lt;/span&gt;&lt;span class=p&gt;:&lt;/span&gt; &lt;span class=n&gt;Conveyance&lt;/span&gt;

&lt;span class=n&gt;ShippingOptions&lt;/span&gt; &lt;span class=o&gt;=&lt;/span&gt; &lt;span class=n&gt;NewType&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=s2&gt;"ShippingOptions"&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt; &lt;span class=n&gt;_RealShipOpts&lt;/span&gt;&lt;span class=p&gt;)&lt;/span&gt;

&lt;span class=k&gt;def&lt;/span&gt; &lt;span class=nf&gt;shipFast&lt;/span&gt;&lt;span class=p&gt;()&lt;/span&gt; &lt;span class=o&gt;-&amp;gt;&lt;/span&gt; &lt;span class=n&gt;ShippingOptions&lt;/span&gt;&lt;span class=p&gt;:&lt;/span&gt;
    &lt;span class=k&gt;return&lt;/span&gt; &lt;span class=n&gt;ShippingOptions&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=n&gt;_RealShipOpts&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=n&gt;Carrier&lt;/span&gt;&lt;span class=o&gt;.&lt;/span&gt;&lt;span class=n&gt;FedEx&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt; &lt;span class=n&gt;Conveyance&lt;/span&gt;&lt;span class=o&gt;.&lt;/span&gt;&lt;span class=n&gt;air&lt;/span&gt;&lt;span class=p&gt;))&lt;/span&gt;

&lt;span class=k&gt;def&lt;/span&gt; &lt;span class=nf&gt;shipNormal&lt;/span&gt;&lt;span class=p&gt;()&lt;/span&gt; &lt;span class=o&gt;-&amp;gt;&lt;/span&gt; &lt;span class=n&gt;ShippingOptions&lt;/span&gt;&lt;span class=p&gt;:&lt;/span&gt;
    &lt;span class=k&gt;return&lt;/span&gt; &lt;span class=n&gt;ShippingOptions&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=n&gt;_RealShipOpts&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=n&gt;Carrier&lt;/span&gt;&lt;span class=o&gt;.&lt;/span&gt;&lt;span class=n&gt;UPS&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt; &lt;span class=n&gt;Conveyance&lt;/span&gt;&lt;span class=o&gt;.&lt;/span&gt;&lt;span class=n&gt;truck&lt;/span&gt;&lt;span class=p&gt;))&lt;/span&gt;

&lt;span class=k&gt;def&lt;/span&gt; &lt;span class=nf&gt;shipSlow&lt;/span&gt;&lt;span class=p&gt;()&lt;/span&gt; &lt;span class=o&gt;-&amp;gt;&lt;/span&gt; &lt;span class=n&gt;ShippingOptions&lt;/span&gt;&lt;span class=p&gt;:&lt;/span&gt;
    &lt;span class=k&gt;return&lt;/span&gt; &lt;span class=n&gt;ShippingOptions&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=n&gt;_RealShipOpts&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=n&gt;Carrier&lt;/span&gt;&lt;span class=o&gt;.&lt;/span&gt;&lt;span class=n&gt;USPS&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt; &lt;span class=n&gt;Conveyance&lt;/span&gt;&lt;span class=o&gt;.&lt;/span&gt;&lt;span class=n&gt;train&lt;/span&gt;&lt;span class=p&gt;))&lt;/span&gt;

&lt;span class=k&gt;def&lt;/span&gt; &lt;span class=nf&gt;shippingDetailed&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;
    &lt;span class=n&gt;carrier&lt;/span&gt;&lt;span class=p&gt;:&lt;/span&gt; &lt;span class=n&gt;Carrier&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt; &lt;span class=n&gt;conveyance&lt;/span&gt;&lt;span class=p&gt;:&lt;/span&gt; &lt;span class=n&gt;Conveyance&lt;/span&gt;
&lt;span class=p&gt;)&lt;/span&gt; &lt;span class=o&gt;-&amp;gt;&lt;/span&gt; &lt;span class=n&gt;ShippingOptions&lt;/span&gt;&lt;span class=p&gt;:&lt;/span&gt;
    &lt;span class=k&gt;return&lt;/span&gt; &lt;span class=n&gt;ShippingOptions&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=n&gt;_RealShipOpts&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=n&gt;carrier&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt; &lt;span class=n&gt;conveyance&lt;/span&gt;&lt;span class=p&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;As a &lt;code&gt;NewType&lt;/code&gt;, our public &lt;code&gt;ShippingOptions&lt;/code&gt; type doesn’t have a constructor.
Since &lt;code&gt;_RealShipOpts&lt;/code&gt; is private, and all its attributes are private, we can
completely remove the old versions.&lt;/p&gt;
&lt;p&gt;Anything &lt;em&gt;within&lt;/em&gt; our shipping library can still access the private variables
on &lt;code&gt;ShippingOptions&lt;/code&gt;; as a &lt;code&gt;NewType&lt;/code&gt;, it’s the same type as its base at
runtime, so it presents minimal&lt;sup id=fnref:1:opaque-types-in-python-2026-5&gt;&lt;a class=footnote-ref href=#fn:1:opaque-types-in-python-2026-5 id=fnref:1&gt;1&lt;/a&gt;&lt;/sup&gt; overhead.&lt;/p&gt;
&lt;p&gt;Clients &lt;em&gt;outside&lt;/em&gt; our shipping library can still call all of our public
constructors: &lt;code&gt;shipFast&lt;/code&gt;, &lt;code&gt;shipNormal&lt;/code&gt;, and &lt;code&gt;shipSlow&lt;/code&gt; all still work with the
same (as far as calling code knows) signature and behavior.&lt;/p&gt;
&lt;p&gt;If you need to build and convey some state within your public API, while
avoiding breakages associated with compatibility churn, hopefully this
technique can help you do that!&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=acknowledgments&gt;Acknowledgments&lt;/h2&gt;
&lt;p class=update-note&gt;Thanks for reading, and thank you to &lt;a href="/pages/patrons.html"&gt;my patrons&lt;/a&gt; who are
supporting my writing on this blog.  If you like what you’ve read here and
you’d like to read more of it, or you’d like to support my &lt;a href="https://github.com/glyph/"&gt;various open-source
endeavors&lt;/a&gt;, you can &lt;a href="/pages/patrons.html"&gt;support my work as a
sponsor&lt;/a&gt;.&lt;/p&gt;
&lt;div class=footnote&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=fn:1:opaque-types-in-python-2026-5&gt;
&lt;p id=fn:1&gt;The overhead is minimal, but it is not &lt;em&gt;completely&lt;/em&gt; zero.  The suggested
idiom for converting to a &lt;code&gt;NewType&lt;/code&gt; is to call it like a function, as I’ve
done in these examples, but if you are wanting to use this pattern &lt;em&gt;inside&lt;/em&gt;
of a hot loop, you can use &lt;code&gt;# type: ignore[return-value]&lt;/code&gt; comments to avoid
that small cost. &lt;a class=footnote-backref href=#fnref:1:opaque-types-in-python-2026-5 title="Jump back to footnote 1 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;&lt;/body&gt;</content><category term="misc"></category><category term="python"></category><category term="programming"></category></entry><entry><title>What Is Code Review For?</title><link href="https://blog.glyph.im/2026/03/what-is-code-review-for.html" rel="alternate"></link><published>2026-03-03T21:24:00-08:00</published><updated>2026-03-03T21:24:00-08:00</updated><author><name>Glyph</name></author><id>tag:blog.glyph.im,2026-03-03:/2026/03/what-is-code-review-for.html</id><summary type="html">&lt;p&gt;Code review is not for catching bugs.&lt;/p&gt;</summary><content type="html">&lt;body&gt;&lt;h2 id=humans-are-bad-at-perceiving&gt;Humans Are Bad At Perceiving&lt;/h2&gt;
&lt;p&gt;Humans are not particularly good at catching bugs.  For one thing, we get tired
easily.  &lt;a href="https://smartbear.com/learn/code-review/best-practices-for-peer-code-review/"&gt;There is some science on this, indicating that humans can’t even
maintain enough concentration to review more than about 400 lines of code at a
time.&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;We have existing terms of art, in various fields, for the ways in which the
human perceptual system fails to register stimuli. Perception fails when humans
are distracted, tired, overloaded, or merely improperly engaged.&lt;/p&gt;
&lt;p&gt;Each of these has implications for the fundamental limitations of code review
as an engineering practice:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href="https://en.wikipedia.org/wiki/Inattentional_blindness"&gt;Inattentional
  Blindness&lt;/a&gt;: you won’t
  be able to reliably find bugs that you’re &lt;em&gt;not&lt;/em&gt; looking for.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href="https://en.wikipedia.org/wiki/Repetition_blindness"&gt;Repetition Blindness&lt;/a&gt;:
  you won’t be able to reliably find bugs that you &lt;em&gt;are&lt;/em&gt; looking for, if they
  keep occurring.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href="https://www.asisonline.org/security-management-magazine/articles/2013/05/vigilance-fatigue/"&gt;Vigilance
  Fatigue&lt;/a&gt;:
  you won’t be able to reliably find either kind of bugs, if you have to keep
  being alert to the presence of bugs all the time.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;and, of course, the distinct but related &lt;a href="https://www.ibm.com/think/topics/alert-fatigue"&gt;Alert
  Fatigue&lt;/a&gt;: you won’t even be
  able to reliably &lt;em&gt;evaluate&lt;/em&gt; reports of &lt;em&gt;possible&lt;/em&gt; bugs, if there are too many
  false positives.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=never-send-a-human-to-do-a-machines-job&gt;Never Send A Human To Do A Machine’s Job&lt;/h2&gt;
&lt;p&gt;When you need to catch a category of error in your code reliably, you will need
a deterministic tool to evaluate — and, thanks to our old friend “alert
fatigue” above — ideally, to also remedy that type of error.  These tools will
relieve the need for a human to make the same repetitive checks over and over.
None of them are perfect, but:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;to catch logical errors, use automated tests.&lt;/li&gt;
&lt;li&gt;to catch formatting errors, use autoformatters.&lt;/li&gt;
&lt;li&gt;to catch common mistakes, use linters.&lt;/li&gt;
&lt;li&gt;to catch common security problems, use a security scanner.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Don’t blame reviewers for missing these things.&lt;/p&gt;
&lt;p&gt;Code review should not be how you catch bugs.&lt;/p&gt;
&lt;h2 id=what-is-code-review-for-then&gt;What Is Code Review For, Then?&lt;/h2&gt;
&lt;p&gt;Code review is for three things.&lt;/p&gt;
&lt;p&gt;First, code review is for catching &lt;em&gt;process failures&lt;/em&gt;.  If a reviewer &lt;em&gt;has&lt;/em&gt;
noticed a few bugs of the same type in code review, that’s a sign that that
type of bug is probably getting &lt;em&gt;through&lt;/em&gt; review more often than it’s getting
caught.  Which means it’s time to figure out a way to deploy a tool or a test
into CI that will &lt;em&gt;reliably&lt;/em&gt; prevent that class of error, without requiring
reviewers to be vigilant to it any more.&lt;/p&gt;
&lt;p&gt;Second — and this is actually its &lt;em&gt;more important&lt;/em&gt; purpose — code review is a
tool for &lt;em&gt;acculturation&lt;/em&gt;.  Even if you already have good tools, good processes,
and good documentation, new members of the team won’t necessarily &lt;em&gt;know&lt;/em&gt; about
those things.  Code review is an opportunity for older members of the team to
introduce newer ones to existing tools, patterns, or areas of responsibility.
If you’re building an observer pattern, you might not realize that the codebase
you’re working in already has an existing idiom for doing that, so you wouldn’t
even think to search for it, but someone else who has worked more with the code
might know about it and help you avoid repetition.&lt;/p&gt;
&lt;p&gt;You will notice that I carefully avoided saying “junior” or “senior” in that
paragraph.  Sometimes the newer team member is actually more senior.  But also,
the acculturation goes both ways.  This is the third thing that code review is
for: &lt;em&gt;disrupting&lt;/em&gt; your team’s culture and avoiding stagnation.  If you have new
talent, a fresh perspective can &lt;em&gt;also&lt;/em&gt; be an extremely valuable tool for
building a healthy culture.  If you’re new to a team and trying to build
something with an observer pattern, and this codebase has no tools for that,
but your &lt;em&gt;last&lt;/em&gt; job did, and it used one from an open source library, that is a
good thing to point out in a review as well.  It’s an opportunity to spot areas
for improvement to culture, as much as it is to spot areas for improvement to
process.&lt;/p&gt;
&lt;p&gt;Thus, code review should be as hierarchically flat as possible.  If the goal of
code review were to spot bugs, it would make sense to reserve the ability to
review code to only the most senior, detail-oriented, rigorous engineers in the
organization.  But most teams already know that that’s a recipe for
brittleness, stagnation and bottlenecks.  Thus, even though we &lt;em&gt;know&lt;/em&gt; that not
everyone on the team will be equally good at spotting bugs, it is very common
in most teams to allow anyone past some fairly low minimum seniority bar to do
reviews, often as low as “everyone on the team who has finished onboarding”.&lt;/p&gt;
&lt;h2 id=oops-surprise-this-post-is-actually-about-llms-again&gt;Oops, Surprise, This Post Is Actually About LLMs Again&lt;/h2&gt;
&lt;p&gt;Sigh.  I’m as disappointed as you are, but there are no two ways about it: LLM
code generators are everywhere now, and we need to talk about how to deal with
them.  Thus, an important corollary of this understanding that code review is a
&lt;em&gt;social activity&lt;/em&gt;, is that LLMs are not &lt;em&gt;social actors&lt;/em&gt;, thus you cannot rely
on code review to inspect their output.&lt;/p&gt;
&lt;p&gt;My own &lt;em&gt;personal&lt;/em&gt; preference would be to eschew their use entirely, but in the
spirit of harm reduction, if you’re going to use LLMs to generate code, you
need to remember the ways in which LLMs are not like human beings.&lt;/p&gt;
&lt;p&gt;When you relate to a human colleague, you will expect that:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;you can make decisions about what to focus on based on their level of
   experience and areas of expertise to know what problems to focus on; from a
   late-career colleague you might be looking for bad habits held over from
   legacy programming languages; from an earlier-career colleague you might be
   focused more on logical test-coverage gaps,&lt;/li&gt;
&lt;li&gt;and, they will learn from repeated interactions so that you can gradually
   focus less on a specific type of problem once you have seen that they’ve
   learned how to address it,&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;With an LLM, by contrast, while errors can certainly be &lt;em&gt;biased&lt;/em&gt; a bit by the
prompt from the engineer and pre-prompts that might exist in the repository, the
types of errors that the LLM will make are somewhat more uniformly distributed
across the experience range.&lt;/p&gt;
&lt;p&gt;You will still find supposedly extremely sophisticated LLMs making &lt;a href="https://arxiv.org/html/2407.07064v2"&gt;extremely
common mistakes&lt;/a&gt;, specifically because
they &lt;em&gt;are&lt;/em&gt; common, and thus appear frequently in the training data.&lt;/p&gt;
&lt;p&gt;The LLM also can’t really learn.  An intuitive response to this problem is to
simply continue adding more and more instructions to its pre-prompt, treating
&lt;em&gt;that&lt;/em&gt; text file as its “memory”, but that &lt;a href="https://github.com/anthropics/claude-code/issues/2766"&gt;just doesn’t work, and probably
never will&lt;/a&gt;.  The
problem — “&lt;a href="https://research.trychroma.com/context-rot"&gt;context rot&lt;/a&gt;” is
somewhat fundamental to the nature of the technology.&lt;/p&gt;
&lt;p&gt;Thus, code-generators must be treated more adversarially than you would a human
code review partner.  When you notice it making errors, you &lt;em&gt;always&lt;/em&gt; have to
add tests to a mechanical, deterministic harness that will evaluates the code,
because the LLM cannot meaningfully learn from its mistakes outside a very
small context window in the way that a human would, so giving it &lt;em&gt;feedback&lt;/em&gt; is
unhelpful.  Asking it to just generate the code again still requires you to
review it &lt;em&gt;all&lt;/em&gt; again, and as we have previously learned, you, a human, cannot
review more than 400 lines at once.&lt;/p&gt;
&lt;h2 id=to-sum-up&gt;To Sum Up&lt;/h2&gt;
&lt;p&gt;Code review is a social process, and you should treat it as such.  When you’re
reviewing code from humans, share knowledge and encouragement as much as you
share bugs or unmet technical requirements.&lt;/p&gt;
&lt;p&gt;If you must reviewing code from an LLM, strengthen your automated code-quality
verification tooling and make sure that its agentic loop will fail on its own
when those quality checks fail immediately next time. Do not fall into the trap
of appealing to its feelings, knowledge, or experience, because it doesn’t have
any of those things.&lt;/p&gt;
&lt;p&gt;But for both humans &lt;em&gt;and&lt;/em&gt; LLMs, do not fall into the trap of thinking that your
code review process is catching your bugs.  That’s not its job.&lt;/p&gt;
&lt;h2 id=acknowledgments&gt;Acknowledgments&lt;/h2&gt;
&lt;p class=update-note&gt;Thank you to &lt;a href="/pages/patrons.html"&gt;my patrons&lt;/a&gt; who are supporting my writing on
this blog.  If you like what you’ve read here and you’d
like to read more of it, or you’d like to support my &lt;a href="https://github.com/glyph/"&gt;various open-source
endeavors&lt;/a&gt;, you can &lt;a href="/pages/patrons.html"&gt;support my work as a
sponsor&lt;/a&gt;!&lt;/p&gt;&lt;/body&gt;</content><category term="misc"></category><category term="programming"></category><category term="process"></category><category term="ai"></category><category term="llm"></category></entry><entry><title>How To Argue With Me About AI, If You Must</title><link href="https://blog.glyph.im/2026/01/how-to-argue-with-me-about-ai.html" rel="alternate"></link><published>2026-01-04T21:22:00-08:00</published><updated>2026-01-04T21:22:00-08:00</updated><author><name>Glyph</name></author><id>tag:blog.glyph.im,2026-01-04:/2026/01/how-to-argue-with-me-about-ai.html</id><summary type="html">&lt;p&gt;If you insist we have a conversation, please come prepared.&lt;/p&gt;</summary><content type="html">&lt;body&gt;&lt;p&gt;As you already know if you’ve read any of this blog in the last few years, I am
a somewhat reluctant — but nevertheless quite staunch — critic of LLMs.  This
means that I have enthusiasts of varying degrees sometimes taking issue with my
stance.&lt;/p&gt;
&lt;p&gt;It seems that I am not going to get away from discussions, and, let’s be
honest, pretty intense arguments about “AI” any time soon.  These arguments are
starting to make me quite upset.  So it might be time to set some rules of
engagement.&lt;/p&gt;
&lt;p&gt;I’ve written about all of these before at greater length, but this is a short
post because it’s not about the technology or making a broader point, it’s
about &lt;em&gt;me&lt;/em&gt;.  These are rules for engaging with me, personally, on this topic.
Others are welcome to adopt these rules if they so wish but I am not
encouraging anyone to do so.&lt;/p&gt;
&lt;p&gt;Thus, I’ve made this post as short as I can so everyone interested in engaging
can read the whole thing.  If you can’t make it through to the end, then please
just follow Rule Zero.&lt;/p&gt;
&lt;h2 id=rule-zero-maybe-dont&gt;Rule Zero: Maybe Don’t&lt;/h2&gt;
&lt;p&gt;You are welcome to ignore me.  You can think my take is stupid and I can think
yours is.  We don’t have to get into an Internet Fight about it; we can even
remain friends.  You do not need to instigate an argument with me at all, if
you think that my analysis is so bad that it doesn’t require rebutting.&lt;/p&gt;
&lt;h2 id=rule-one-no-just&gt;Rule One: No ‘Just’&lt;/h2&gt;
&lt;p&gt;As I explained in a post with perhaps the least-predictive title I’ve ever
written, &lt;a href="https://blog.glyph.im/2025/06/i-think-im-done-thinking-about-genai-for-now.html"&gt;“I Think I’m Done Thinking About genAI For
Now”&lt;/a&gt;, I’ve already
heard a bunch of bad arguments.  Don’t tell me to ‘just’ use a better model,
use an agentic tool, use a more recent version, or use some prompting trick
that you personally believe works better.  If you skim my work and think that I
must not have deeply researched anything or read about it because you don’t
like my conclusion, that is wrong.&lt;/p&gt;
&lt;h2 id=rule-two-no-look-at-this-cool-thing&gt;Rule Two: No ‘Look At This Cool Thing’&lt;/h2&gt;
&lt;p&gt;Purely as a productivity tool, I have had a terrible experience with genAI.
Perhaps you have had a great one.  Neat.  That’s great for you.  As I explained
&lt;em&gt;at great length&lt;/em&gt; in &lt;a href="https://blog.glyph.im/2025/08/futzing-fraction.html"&gt;“The Futzing Fraction”&lt;/a&gt;,
my concern with generative AI is that &lt;em&gt;I believe&lt;/em&gt; it is probably a &lt;em&gt;net
negative&lt;/em&gt; impact on productivity, based on both my experience and plenty of
citations. Go check out the copious footnotes if you’re interested in more
detail.&lt;/p&gt;
&lt;p&gt;Therefore, I have already acknowledged that you can get an LLM to do various
impressive, cool things, &lt;em&gt;sometimes&lt;/em&gt;.  If I tell you that you will, on average,
lose money betting on a slot machine, &lt;em&gt;a picture of a slot machine hitting a
jackpot is not evidence against my position&lt;/em&gt;.&lt;/p&gt;
&lt;h3 id=rule-two-and-a-half-engage-in-metacognition&gt;Rule Two And A Half: Engage In Metacognition&lt;/h3&gt;
&lt;p&gt;I specifically didn’t title the previous rule “no anecdotes” because data
beyond anecdotes may be extremely expensive to produce.  I don’t want to say
you can never talk to me unless you’re doing a randomized controlled trial.
However, if you are going to tell me an anecdote about the way that you’re
using an LLM, I am interested in hearing &lt;em&gt;how you are compensating&lt;/em&gt; for the
well-documented biases that LLM use tends to induce.  Try to measure what you
can.&lt;/p&gt;
&lt;h2 id=rule-three-do-not-cite-the-deep-magic-to-me&gt;Rule Three: Do Not Cite The Deep Magic To Me&lt;/h2&gt;
&lt;p&gt;As I explained in &lt;a href="https://blog.glyph.im/2024/05/grand-unified-ai-hype.html"&gt;“A Grand Unified Theory of the AI Hype
Cycle”&lt;/a&gt;, I already know &lt;em&gt;quite a bit&lt;/em&gt; of
history of the “AI” label.  If you are tempted to tell me something about how
“AI” is really such a broad field, and it doesn’t just mean LLMs, especially if
you are &lt;em&gt;trying&lt;/em&gt; to launder the reputation of LLMs under the banner of jumbling
them together with other things that have been called “AI”, I assure you that
this will not be convincing to me.&lt;/p&gt;
&lt;h2 id=rule-four-ethics-are-not-optional&gt;Rule Four: Ethics Are Not Optional&lt;/h2&gt;
&lt;p&gt;I have made several arguments in my previous writing: there are ethical
arguments, efficacy arguments, structuralist arguments, efficiency arguments
and aesthetic arguments.&lt;/p&gt;
&lt;p&gt;I am happy to, for the purposes of a good-faith discussion, focus on a specific
set of concerns or an individual point that you want to make where you think I
got something wrong.  If you convince me that I am entirely incorrect about the
effectiveness or predictability of LLMs in general or as specific LLM product,
you don’t need to make a comprehensive argument about whether one should use
the technology overall.  I will even assume that you &lt;em&gt;have&lt;/em&gt; your own ethical
arguments.&lt;/p&gt;
&lt;p&gt;However, if you scoff at the idea that one &lt;em&gt;should have any ethical boundaries
at all&lt;/em&gt;, and think that there’s no reason to care about the overall utilitarian
impact of this technology, that it’s worth using no matter what else it does as
long as it makes you 5% better at your job, that’s sociopath behavior.&lt;/p&gt;
&lt;p&gt;This includes extreme whataboutism regarding things like the water use of
datacenters, other elements of the surveillance technology stack, and so on.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=consequences&gt;Consequences&lt;/h2&gt;
&lt;p&gt;These are rules, once again, just for engaging with &lt;em&gt;me&lt;/em&gt;. I have no particular
power to enact broader sanctions upon you, nor would I be inclined to do so if
I could.  However, if you can’t stay within these basic parameters &lt;em&gt;and&lt;/em&gt; you
insist upon continuing to direct messages to me about this topic, I will
summarily block you with no warning, on mastodon, email, GitHub, IRC, or
wherever else you’re choosing to do that.  This is for your benefit as well:
such a discussion will not be a productive use of either of our time.&lt;/p&gt;&lt;/body&gt;</content><category term="misc"></category><category term="ai"></category><category term="meta"></category></entry><entry><title>The Next Thing Will Not Be Big</title><link href="https://blog.glyph.im/2026/01/the-next-thing-will-not-be-big.html" rel="alternate"></link><published>2026-01-01T17:59:00-08:00</published><updated>2026-01-01T17:59:00-08:00</updated><author><name>Glyph</name></author><id>tag:blog.glyph.im,2026-01-01:/2026/01/the-next-thing-will-not-be-big.html</id><summary type="html">&lt;p&gt;Disruption, too, will be disrupted.&lt;/p&gt;</summary><content type="html">&lt;body&gt;&lt;p&gt;The dawning of a new year is an opportune moment to contemplate what has
transpired in the old year, and consider what is likely to happen in the new
one.&lt;/p&gt;
&lt;p&gt;Today, I’d like to contemplate that contemplation itself.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;The 20th century was an era characterized by rapidly accelerating change in
technology and industry, creating shorter and shorter cultural cycles of
changes in lifestyles. Thus far, the 21st century seems to be following that
trend, at least in its recently concluded first quarter.&lt;/p&gt;
&lt;p&gt;The early half of the twentieth century saw the massive disruption caused by
electrification, radio, motion pictures, and then television.&lt;/p&gt;
&lt;p&gt;In 1971, Intel poured gasoline on that fire by releasing the 4004, a microchip
generally recognized as the first general-purpose microprocessor. Popular
innovations rapidly followed: the computerized cash register, the personal
computer, credit cards, cellular phones, text messaging, the Internet, the web,
online games, mass surveillance, app stores, social media.&lt;/p&gt;
&lt;p&gt;These innovations have arrived faster than previous generations, but also, they
have crossed a crucial threshold: that of the human lifespan.&lt;/p&gt;
&lt;p&gt;While the entire second millennium A.D. has been characterized by a gradually
accelerating rate of technological and social change — the printing press and
the industrial revolution were no slouches, in terms of changing society, and
those predate the 20th century — most of those changes had the benefit of
unfolding throughout the course of a generation or so.&lt;/p&gt;
&lt;p&gt;Which means that any &lt;em&gt;individual person&lt;/em&gt; in any given century up to the 20th
might remember &lt;em&gt;one&lt;/em&gt; major world-altering social shift within their lifetime,
not five to ten of them.  The diversity of human experience is vast, but &lt;em&gt;most&lt;/em&gt;
people would not &lt;em&gt;expect&lt;/em&gt; that the defining technology of their lifetime was
merely the latest in a progression of predictable civilization-shattering
marvels.&lt;/p&gt;
&lt;p&gt;Along with each of these successive generations of technology, we minted a new
generation of industry titans. Westinghouse, Carnegie, Sarnoff, Edison, Ford,
Hughes, Gates, Jobs, Zuckerberg, Musk. Not just individual rich people, but
entire new &lt;em&gt;classes&lt;/em&gt; of rich people that did not exist before. “Radio DJ”,
“Movie Star”, “Rock Star”, “Dot Com Founder”, were all new paths to wealth
opened (and closed) by specific technologies. While most of these people did
come from at least &lt;em&gt;some&lt;/em&gt; level of generational wealth, they no longer came
from a literal hereditary aristocracy.&lt;/p&gt;
&lt;p&gt;To &lt;em&gt;describe&lt;/em&gt; this new feeling of constant acceleration, a new phrase was
coined: “&lt;a href="https://grammarphobia.com/blog/2015/11/thing-2.html"&gt;The Next Big
Thing&lt;/a&gt;”.  In addition to
denoting that some Thing was coming and that it would be Big (i.e.: that it
would change a lot about our lives), this phrase also carries the strong
&lt;em&gt;implication&lt;/em&gt; that such a Thing would be a product.  Not a development in
social relationships or a shift in cultural values, but some new and amazing
form of conveying salted &lt;a href="https://en.wikipedia.org/wiki/Spam_(food)"&gt;meat&lt;/a&gt;
&lt;a href="https://en.wikipedia.org/wiki/Bovril"&gt;paste&lt;/a&gt; or what-have-you, that would make
whatever lucky tinkerer who stumbled into it into a billionaire — along with
any friends and family lucky enough to believe in their vision and get in on
the ground floor with an investment.&lt;/p&gt;
&lt;p&gt;In the latter part of the 20th century, our entire model of capital allocation
shifted to account for this widespread belief. No longer were mega-businesses
built by bank loans, stock issuances, and reinvestment of profit, the new model
was “Venture Capital”. Venture capital is a model of capital allocation
&lt;em&gt;explicitly predicated&lt;/em&gt; on the idea that carefully considering each bet on a
likely-to-succeed business and reducing one’s risk was a waste of time, because
the return on the equity from the Next Big Thing would be so disproportionately
huge — 10x, 100x, 1000x – that one could afford to make &lt;em&gt;at least&lt;/em&gt; 10 bad bets
for each good one, and still come out ahead.&lt;/p&gt;
&lt;p&gt;The biggest risk was in &lt;em&gt;missing the deal&lt;/em&gt;, not in giving a bunch of money to a
scam.  Thus, value investing and focus on fundamentals have been broadly
disregarded in favor of the pursuit of the Next Big Thing.&lt;/p&gt;
&lt;p&gt;If Americans of the twentieth century were temporarily embarrassed
millionaires, those of the twenty-first are all temporarily embarrassed
&lt;a href="https://en.wikipedia.org/wiki/Big_Tech#Acronyms"&gt;FAANG&lt;/a&gt; CEOs.&lt;/p&gt;
&lt;p&gt;The predicament that this tendency leaves us in today is that the world is
increasingly run by generations — GenX and Millennials — with the shared
experience that the computer industry, either hardware or software, would
produce some radical innovation every few years.  We assume that to be true.&lt;/p&gt;
&lt;p&gt;But all things change, even change itself, and that industry is beginning to
slow down.  Physically, transistor density is starting to &lt;a href="https://interestingengineering.com/innovation/transistors-moores-law"&gt;brush up against
physical
limits&lt;/a&gt;.
Economically, most people are drowning in more compute power than they know
what to do with anyway. Users already have most of what they need from the
Internet.&lt;/p&gt;
&lt;p&gt;The big new feature in every operating system is a bunch of &lt;a href="https://www.cnet.com/tech/mobile/73-of-iphone-owners-say-no-thanks-to-apple-intelligence-new-data-echoes-cnets-findings/"&gt;useless
junk&lt;/a&gt;
&lt;a href="https://www.windowscentral.com/microsoft/windows-11/2025-has-been-an-awful-year-for-windows-11-with-infuriating-bugs-and-constant-unwanted-features"&gt;nobody really
wants&lt;/a&gt;
and is seeing remarkably little uptake.  Social media and smartphones changed
the world, true, but… those are both innovations from 2008.  They’re just not
&lt;em&gt;new&lt;/em&gt; any more.&lt;/p&gt;
&lt;p&gt;So we are all — collectively, culturally — looking for the Next Big Thing, and
we keep not finding it.&lt;/p&gt;
&lt;p&gt;It wasn’t 3D printing. It wasn’t crowdfunding. It wasn’t smart watches. It
wasn’t VR. It wasn’t the Metaverse, it wasn’t Bitcoin, it wasn’t NFTs&lt;sup id=fnref:1:the-next-thing-will-not-be-big-2026-1&gt;&lt;a class=footnote-ref href=#fn:1:the-next-thing-will-not-be-big-2026-1 id=fnref:1&gt;1&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;It’s also not AI, but this is why so many people &lt;em&gt;assume&lt;/em&gt; that it will be AI.
Because it’s got to be &lt;em&gt;something&lt;/em&gt;, right?  If it’s got to be &lt;em&gt;something&lt;/em&gt; then
AI is as good a guess as anything else right now.&lt;/p&gt;
&lt;p&gt;The fact is, &lt;em&gt;our lifetimes have been an extreme anomaly&lt;/em&gt;.  Things like the
Internet used to come along every thousand years or so, and while we might
expect that the pace will stay a bit higher than that, it is not reasonable to
expect that something new &lt;em&gt;like&lt;/em&gt; “personal computers” or “the Internet”&lt;sup id=fnref:3:the-next-thing-will-not-be-big-2026-1&gt;&lt;a class=footnote-ref href=#fn:3:the-next-thing-will-not-be-big-2026-1 id=fnref:3&gt;3&lt;/a&gt;&lt;/sup&gt;
will arrive again.&lt;/p&gt;
&lt;p&gt;We are not going to get rich by getting in on the ground floor of the next
Apple or the next Google because the next Apple and the next Google are Apple
and Google.  The industry is maturing.  Software technology, computer
technology, and internet technology are all maturing.&lt;/p&gt;
&lt;h2 id=there-will-be-next-things&gt;There Will Be Next Things&lt;/h2&gt;
&lt;p&gt;Research and development is happening in all fields all the time. Amazing new
developments quietly and regularly occur in pharmaceuticals and in materials
science.  But these are not predictable. They do not inhabit the public
consciousness until they’ve already happened, and they are rarely so profound
and transformative that they change &lt;em&gt;everybody’s&lt;/em&gt; life.&lt;/p&gt;
&lt;p&gt;There will even be new things in the computer industry, both software and
hardware. Foldable phones do address a real problem (I wish the screen were
even bigger but I don’t want to carry around such a big device), and would
probably be more popular if they got the costs under control.  One day
somebody’s going to crack the problem of volumetric displays, probably. Some VR
product will probably, eventually, hit a more realistic price/performance ratio
where the niche will expand at least a little more.&lt;/p&gt;
&lt;p&gt;Maybe there will even be something genuinely useful, which is recognizably
adjacent to the current “AI” fad, but if it is, it will be some &lt;em&gt;new
development&lt;/em&gt; that we haven’t seen yet.  If current AI technology were
sufficient to drive some interesting product, it would already be doing it, not
&lt;a href="https://theoutpost.ai/news-story/major-study-reveals-ai-benchmarks-may-be-misleading-casting-doubt-on-reported-capabilities-21513/"&gt;using marketing disguised as
science&lt;/a&gt;
to &lt;a href="https://www.wired.com/story/the-ai-industrys-scaling-obsession-is-headed-for-a-cliff/"&gt;conceal diminishing
returns&lt;/a&gt;
on current investments.&lt;/p&gt;
&lt;h2 id=but-they-will-not-be-big&gt;But They Will Not Be Big&lt;/h2&gt;
&lt;p&gt;The impulse to find the One Big Thing that will dominate the next five years is
a fool’s errand.  Incremental gains are diminishing across the board.  The
markets for time and attention&lt;sup id=fnref:2:the-next-thing-will-not-be-big-2026-1&gt;&lt;a class=footnote-ref href=#fn:2:the-next-thing-will-not-be-big-2026-1 id=fnref:2&gt;2&lt;/a&gt;&lt;/sup&gt; are largely saturated.  There’s no need for
another streaming service if 100% of your leisure time is already committed to
TikTok, YouTube and Netflix; famously, Netflix has already considered
&lt;a href="https://www.fastcompany.com/40491939/netflix-ceo-reed-hastings-sleep-is-our-competition"&gt;sleep&lt;/a&gt;
its primary competitor for close to a decade - years &lt;em&gt;before&lt;/em&gt; the pandemic.&lt;/p&gt;
&lt;p&gt;Those rare tech markets which &lt;em&gt;aren’t&lt;/em&gt; saturated are suffering from pedestrian
economic problems like wealth inequality, not technological bottlenecks.&lt;/p&gt;
&lt;p&gt;For example, the thing preventing the development of a robot that can do your
laundry and your dishes without your input is not necessarily that we couldn’t
build something like that, but that most households just &lt;em&gt;can’t afford it&lt;/em&gt;
without &lt;a href="https://www.epi.org/productivity-pay-gap/"&gt;wage growth catching up to productivity
growth&lt;/a&gt;.  It doesn’t make sense for
anyone to commit to the substantial R&amp;amp;D investment that such a thing would
take, if the market doesn’t exist because the average worker isn’t paid enough
to afford it on top of all the &lt;em&gt;other&lt;/em&gt; tech which is already required to exist
in society.&lt;/p&gt;
&lt;p&gt;The projected income from the tiny, wealthy sliver of the population who
&lt;em&gt;could&lt;/em&gt; pay for the hardware, cannot justify an investment in the software past
a &lt;a href="https://futurism.com/future-society/robot-servant-neo-remote-controlled"&gt;fake version remotely operated by workers in the global south, only made
possible by Internet wage
arbitrage&lt;/a&gt;,
i.e. a more palatable, modern version of indentured servitude.&lt;/p&gt;
&lt;p&gt;Even if we were to accept the premise of an actually-“AI” version of this, that
is still just a wish that ChatGPT could somehow improve enough behind the
scenes to replace that worker, not any substantive investment in a novel,
proprietary-to-the-chores-robot software system which could &lt;em&gt;reliably&lt;/em&gt; perform
specific functions.&lt;/p&gt;
&lt;h2 id=what-then&gt;What, Then?&lt;/h2&gt;
&lt;p&gt;The expectation for, and lack of, a “big thing” is a big problem.  There are
others who could describe its economic, political, and financial dimensions
better than I can.  So then let me speak to my expertise and my audience: open
source software developers.&lt;/p&gt;
&lt;p&gt;When I began my own involvement with open source, a big part of the draw for me
was participating in a low-cost (to the corporate developer) but high-value (to
society at large) positive externality.  None of my employers would ever have
cared about many of the &lt;a href="https://deluge-torrent.org"&gt;applications&lt;/a&gt; for which
&lt;a href="https://twisted.org/"&gt;Twisted&lt;/a&gt; forms a core bit of infrastructure; nor would I
have been able to predict those applications’ existence.  Yet, it is nice to
have contributed to their development, even a little bit.&lt;/p&gt;
&lt;p&gt;However, it’s not actually a positive externality if the public at large can’t
directly &lt;em&gt;benefit&lt;/em&gt; from it.&lt;/p&gt;
&lt;p&gt;When &lt;em&gt;real&lt;/em&gt; world-changing, disruptive developments are occurring, the
bean-counters are not watching positive externalities too closely.  As we
discovered with &lt;a href="https://www.businessinsider.com/zirp-end-of-cushy-big-tech-job-perks-mass-layoffs-2024-2"&gt;many of the other benefits that temporarily accrued to
labor&lt;/a&gt;
in the tech economy, Open Source that is &lt;em&gt;usable by individuals and small
companies&lt;/em&gt; may have been a ZIRP.  If you know you’re gonna make a billion
dollars you’re not going to worry about giving away a few hundred thousand here
and there.&lt;/p&gt;
&lt;p&gt;When gains are smaller and harder to realize, and margins are starting to get
squeezed, it’s harder to justify the investment in vaguely good vibes.&lt;/p&gt;
&lt;p&gt;But this, itself, is not a call to action.  I doubt very much that anyone
reading this can do anything about the macroeconomic reality of higher interest
rates. The technological reality of “development is happening slower” is
inherently something that you can’t change on purpose.&lt;/p&gt;
&lt;p&gt;However, what we &lt;em&gt;can&lt;/em&gt; do is to be aware of this trend in our own work.&lt;/p&gt;
&lt;h2 id=fight-scale-creep&gt;Fight Scale Creep&lt;/h2&gt;
&lt;p&gt;It seems to me that more and more open source infrastructure projects are tools
for hyper-scale application development, only relevant to massive cloud
companies.  This is just a subjective assessment on my part — I’m not sure what
tools even exist today to measure this empirically — but I remember a big part
of the open source community when I was younger being things like Inkscape,
Themes.Org and Slashdot, not React, Docker Hub and Hacker News.&lt;/p&gt;
&lt;p&gt;This is not to say that the hobbyist world no longer exists. There is of course
a ton of stuff going on with Raspberry Pi, Home Assistant, OwnCloud, and so on.
If anything there’s a bit of a resurgence of self-hosting.  But the interests
of self-hosters and corporate developers are growing apart; there seems to be
far less of a beneficial overflow from corporate infrastructure projects into
these enthusiast or prosumer communities.&lt;/p&gt;
&lt;p&gt;This is the concrete call to action: if you are employed in any capacity as an
open source maintainer, dedicate &lt;em&gt;more&lt;/em&gt; energy to medium- or small-scale open
source projects.&lt;/p&gt;
&lt;p&gt;If your assumption is that you will eventually reach a hyper-scale inflection
point, then mimicking Facebook and Netflix is likely to be a good idea.
However, if we can all admit to ourselves that we’re &lt;em&gt;not&lt;/em&gt; going to achieve a
trillion-dollar valuation and a hundred thousand engineer headcount, we can
begin to consider ways to make our Next Thing a bit smaller, and to accommodate
the world as it is rather than as we wish it would be.&lt;/p&gt;
&lt;h2 id=be-prepared-to-scale-down&gt;Be Prepared to Scale Down&lt;/h2&gt;
&lt;p&gt;Here are some design guidelines you might consider, for just about any open
source project, particularly infrastructure ones:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Don’t assume that your software can sustain an arbitrarily large fixed
   overhead because “you just pay that cost once” and you’re going to be
   running a billion instances so it will always amortize; maybe you’re only
   going to be running ten.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Remember that such fixed overhead includes not just CPU, RAM, and filesystem
   storage, but also the learning curve for developers.  Front-loading a
   massive amount of conceptual complexity to accommodate the problems of
   hyper-scalers is a common mistake.  Try to smooth out these complexities and
   introduce them only when necessary.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Test your code on edge devices. This means supporting Windows and macOS, and
   even Android and iOS.  If you want your tool to help empower individual
   users, you will need to meet them where they are, which is &lt;em&gt;not&lt;/em&gt; on an EC2
   instance.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;This includes considering Desktop Linux as a platform, as opposed to Server
   Linux as a platform, which (while they certainly have plenty in common) they
   are also distinct in some details.  Consider the highly specific example of
   secret storage: if you are writing something that intends to live in a cloud
   environment, and you need to configure it with a secret, you will probably
   want to provide it via a text file or an environment variable.  By contrast,
   if you want this same code to run on a desktop system, your users will
   expect you to support the &lt;a href="https://specifications.freedesktop.org/secret-service/latest/"&gt;Secret
   Service&lt;/a&gt;.
   This will likely only require a few lines of code to accommodate, but it is
   a massive difference to the user experience.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Don’t rely on LLMs remaining cheap or free.  If you have LLM-related
   features&lt;sup id=fnref:5:the-next-thing-will-not-be-big-2026-1&gt;&lt;a class=footnote-ref href=#fn:5:the-next-thing-will-not-be-big-2026-1 id=fnref:5&gt;4&lt;/a&gt;&lt;/sup&gt;, make sure that they are sufficiently severable from the rest of
   your offering that if ChatGPT starts costing $1000 a month, your tool
   doesn’t break completely.  Similarly, do not require that your users have
   easy access to half a terabyte of VRAM and a rack full of 5090s in order to
   run a local model.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Even if you &lt;em&gt;were&lt;/em&gt; going to scale up to infinity, the ability to scale down and
consider smaller deployments means that you can run more comfortably on, for
example, a developer’s laptop.  So even if you can’t convince your employer
that this is where the economy and the future of technology in our lifetimes is
going, it can be easy enough to justify this sort of design shift, particularly
as individual choices.  Make your onboarding cheaper, your development feedback loops tighter, and your systems generally more resilient to economic headwinds.&lt;/p&gt;
&lt;p&gt;So, please design your open source libraries, applications, and services to run
on smaller devices, with less complexity.  It will be worth your time as well
as your users’.&lt;/p&gt;
&lt;p&gt;But if you &lt;em&gt;can&lt;/em&gt; fix the whole wealth inequality thing, do that first.&lt;/p&gt;
&lt;h2 id=acknowledgments&gt;Acknowledgments&lt;/h2&gt;
&lt;p class=update-note&gt;Thank you to &lt;a href="/pages/patrons.html"&gt;my patrons&lt;/a&gt; who are supporting my writing on
this blog.  If you like what you’ve read here and you’d
like to read more of it, or you’d like to support my &lt;a href="https://github.com/glyph/"&gt;various open-source
endeavors&lt;/a&gt;, you can &lt;a href="/pages/patrons.html"&gt;support my work as a
sponsor&lt;/a&gt;!&lt;/p&gt;
&lt;div class=footnote&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=fn:1:the-next-thing-will-not-be-big-2026-1&gt;
&lt;p id=fn:1&gt;&lt;a href="https://www.technologyreview.com/10-breakthrough-technologies/2013/"&gt;These sorts of
lists&lt;/a&gt; are
pretty funny reads, in retrospect. &lt;a class=footnote-backref href=#fnref:1:the-next-thing-will-not-be-big-2026-1 title="Jump back to footnote 1 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:2:the-next-thing-will-not-be-big-2026-1&gt;
&lt;p id=fn:2&gt;Which is to say, “distraction”. &lt;a class=footnote-backref href=#fnref:2:the-next-thing-will-not-be-big-2026-1 title="Jump back to footnote 2 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:3:the-next-thing-will-not-be-big-2026-1&gt;
&lt;p id=fn:3&gt;... or even their lesser-but-still-profound aftershocks like “Social
Media”, “Smartphones”, or “On-Demand Streaming Video” ...
secondary manifestations of the underlying innovation of a packet-switched
global digital network ... &lt;a class=footnote-backref href=#fnref:3:the-next-thing-will-not-be-big-2026-1 title="Jump back to footnote 3 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:5:the-next-thing-will-not-be-big-2026-1&gt;
&lt;p id=fn:5&gt;My preference would of course be that you just didn’t have such features
at all, but perhaps even if you agree with me, you are part of an
organization with some mandate to implement LLM stuff.  Just try not to
wrap the chain of this anchor &lt;em&gt;all&lt;/em&gt; the way around your code’s neck. &lt;a class=footnote-backref href=#fnref:5:the-next-thing-will-not-be-big-2026-1 title="Jump back to footnote 4 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;&lt;/body&gt;</content><category term="misc"></category><category term="startups"></category><category term="ai"></category><category term="open-source"></category><category term="programming"></category><category term="politics"></category></entry><entry><title>The “Dependency Cutout” Workflow Pattern, Part I</title><link href="https://blog.glyph.im/2025/11/dependency-cutout-workflow-pattern.html" rel="alternate"></link><published>2025-11-10T17:44:00-08:00</published><updated>2025-11-10T17:44:00-08:00</updated><author><name>Glyph</name></author><id>tag:blog.glyph.im,2025-11-10:/2025/11/dependency-cutout-workflow-pattern.html</id><summary type="html">&lt;p&gt;It’s important to be able to fix bugs in your open source
dependencies, and not just work around them.&lt;/p&gt;</summary><content type="html">&lt;body&gt;&lt;p&gt;Tell me if you’ve heard this one before.&lt;/p&gt;
&lt;p&gt;You’re working on an application.  Let’s call it “FooApp”.  FooApp has a
dependency on an open source library, let’s call it “LibBar”.  You find a bug
in LibBar that affects FooApp.&lt;/p&gt;
&lt;p&gt;To envisage the best possible version of this scenario, let’s say you actively
&lt;em&gt;like&lt;/em&gt; LibBar, both technically and socially.  You’ve contributed to it in the
past.  But this bug is causing production issues in FooApp &lt;em&gt;today&lt;/em&gt;, and
LibBar’s release schedule is quarterly.  FooApp is your job; LibBar is (at
best) your hobby.  Blocking on the full upstream contribution cycle and waiting
for a release is an absolute non-starter.&lt;/p&gt;
&lt;p&gt;What do you do?&lt;/p&gt;
&lt;p&gt;There are a few common reactions to this type of scenario, all of which are
bad options.&lt;/p&gt;
&lt;p&gt;I will enumerate them specifically here, because I suspect that some of them
may resonate with many readers:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Find an alternative to LibBar, and switch to it.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is a bad idea because a transition to a core infrastructure component
could be extremely expensive.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Vendor LibBar into your codebase and fix your vendored version.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is a bad idea because carrying this one fix now requires you to
maintain all the tooling associated with a monorepo&lt;sup id=fnref:1:dependency-cutout-workflow-pattern-2025-11&gt;&lt;a class=footnote-ref href=#fn:1:dependency-cutout-workflow-pattern-2025-11 id=fnref:1&gt;1&lt;/a&gt;&lt;/sup&gt;: you have to be
able to start pulling in new versions from LibBar regularly, reconcile your
changes even though you now have a separate version history on your
imported version, and so on.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://en.wikipedia.org/wiki/Monkey_patch"&gt;Monkey-patch&lt;/a&gt; LibBar to
   include your fix.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is a bad idea because you are now extremely tightly coupled to a
specific version of LibBar.  By modifying LibBar internally like this,
you’re inherently violating its compatibility contract, in a way which is
going to be extremely difficult to test.  You &lt;em&gt;can&lt;/em&gt; test this change, of
course, but as LibBar changes, you will need to replicate any relevant
portions of its test suite (which may be its &lt;em&gt;entire&lt;/em&gt; test suite) in
FooApp.  Lots of potential duplication of effort there.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Implement a workaround in your own code, rather than fixing it.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is a bad idea because you are distorting the responsibility for
correct behavior.  LibBar is supposed to do LibBar’s job, and unless you
have a full wrapper for it in your own codebase, other engineers (including
“yourself, personally”) might later forget to go through the alternate,
workaround codepath, and invoke the buggy LibBar behavior again in some new
place.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Implement the fix upstream in LibBar anyway, because that’s the Right
   Thing To Do, and burn credibility with management while you anxiously wait
   for a release with the bug in production.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is a bad idea because you are betraying your users — by allowing the
buggy behavior to persist — for the workflow convenience of your dependency
providers. Your users are probably giving you money, and trusting you with
their data. This means you have both ethical and economic obligations to
consider their interests.&lt;/p&gt;
&lt;p&gt;As much as it’s nice to participate in the open source community and take
on an appropriate level of burden to maintain the commons, this cannot
sustainably be at the explicit expense of the population you serve
directly.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Even if&lt;/em&gt; we only care about the open source maintainers here, there’s
still a problem: as you are likely to come under immediate pressure to ship
your changes, you will inevitably relay at least a bit of that stress to
the maintainers.  Even if you try to be exceedingly polite, the maintainers
will know that &lt;em&gt;you&lt;/em&gt; are coming under fire for not having shipped the fix
yet, and are likely to feel an even greater burden of obligation to ship
your code fast.&lt;/p&gt;
&lt;p&gt;Much as it’s good to contribute the fix, it’s not great to put this on the
maintainers.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The respective incentive structures of software development — specifically, of
corporate application development and open source infrastructure development —
make options 1-4 very common.&lt;/p&gt;
&lt;p&gt;On the corporate / application side, these issues are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;it’s difficult for corporate developers to get clearance to spend even small amounts of
  their work hours on upstream open source projects, but clearance to spend
  time on the project they actually work on is implicit.  If it takes 3 hours
  of wrangling with Legal&lt;sup id=fnref:2:dependency-cutout-workflow-pattern-2025-11&gt;&lt;a class=footnote-ref href=#fn:2:dependency-cutout-workflow-pattern-2025-11 id=fnref:2&gt;2&lt;/a&gt;&lt;/sup&gt; and 3 hours of implementation work to fix the
  issue in LibBar, but 0 hours of wrangling with Legal and 40 hours of
  implementation work in FooApp, a FooApp developer will often perceive it as
  “easier” to fix the issue downstream.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;it’s difficult for corporate developers to get clearance from management to
  spend even small amounts of &lt;em&gt;money&lt;/em&gt; sponsoring upstream reviewers, so even if
  they can find the time to contribute the fix, chances are high that it will
  remain stuck in review unless they are personally well-integrated members of
  the LibBar development team already.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;even assuming there’s zero pressure whatsoever to avoid open sourcing the
  upstream changes, there’s still the fact inherent to any development team
  that FooApp’s developers will be more familiar with FooApp’s codebase and
  development processes than they are with LibBar’s.  It’s just &lt;em&gt;easier&lt;/em&gt; to
  work there, even if all other things are equal.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;systems for tracking risk from open source dependencies often lack visibility
  into vendoring, particularly if you’re doing a hybrid approach and only
  vendoring a &lt;em&gt;few&lt;/em&gt; things to address work in progress, rather than a
  comprehensive and disciplined approach to a monorepo.  If you fully absorb a
  vendored dependency and then modify it, Dependabot isn’t going to tell you
  that a new version is available any more, because it won’t be present in your
  dependency list.  Organizationally this is bad of course but from the
  perspective of an &lt;em&gt;individual developer&lt;/em&gt; this manifests mostly as fewer
  annoying emails.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But there are problems on the open source side as well.  Those problems are all
derived from one big issue: because we’re often working with relatively small
sums of money, it’s hard for upstream open source developers to &lt;em&gt;consume&lt;/em&gt;
either money or patches from application developers.  It’s nice to say that you
should contribute money to your dependencies, and you absolutely &lt;em&gt;should&lt;/em&gt;, but
the cost-benefit function is discontinuous.  Before a project reaches the
fiscal threshold where it can be at least &lt;em&gt;one&lt;/em&gt; person’s full-time job to worry
about this stuff, there’s often no-one responsible in the first place.
Developers will therefore gravitate to the issues that are either fun, or
relevant to their &lt;em&gt;own&lt;/em&gt; job.&lt;/p&gt;
&lt;p&gt;These mutually-reinforcing incentive structures are a big reason that users of
open source infrastructure, even teams who work at corporate users with
zillions of dollars, don’t reliably contribute back.&lt;/p&gt;
&lt;h2 id=the-answer-we-want&gt;The Answer We Want&lt;/h2&gt;
&lt;p&gt;All those options are bad. If we had a good option, what would it look like?&lt;/p&gt;
&lt;p&gt;It is both practically necessary&lt;sup id=fnref:3:dependency-cutout-workflow-pattern-2025-11&gt;&lt;a class=footnote-ref href=#fn:3:dependency-cutout-workflow-pattern-2025-11 id=fnref:3&gt;3&lt;/a&gt;&lt;/sup&gt; and morally required&lt;sup id=fnref:4:dependency-cutout-workflow-pattern-2025-11&gt;&lt;a class=footnote-ref href=#fn:4:dependency-cutout-workflow-pattern-2025-11 id=fnref:4&gt;4&lt;/a&gt;&lt;/sup&gt; for you to have a
way to temporarily rely on a modified version of an open source dependency,
&lt;em&gt;without&lt;/em&gt; permanently diverging.&lt;/p&gt;
&lt;p&gt;Below, I will describe a desirable abstract workflow for achieving this goal.&lt;/p&gt;
&lt;h3 id=step-0-report-the-problem&gt;Step 0: Report the Problem&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;Before&lt;/em&gt; you get started with any of these other steps, write up a clear
description of the problem and report it to the project as an issue;
specifically, &lt;em&gt;in contrast to&lt;/em&gt; writing it up as a pull request.  Describe the
problem &lt;em&gt;before&lt;/em&gt; submitting a solution.&lt;/p&gt;
&lt;p&gt;You may not be able to wait for a volunteer-run open source project to respond
to your request, but you should &lt;em&gt;at least&lt;/em&gt; tell the project what you’re
planning on doing.&lt;/p&gt;
&lt;p&gt;If you don’t hear back from them at all, you will have at least made sure to
comprehensively describe your issue and strategy beforehand, which will provide
some clarity and focus to your changes.&lt;/p&gt;
&lt;p&gt;If you &lt;em&gt;do&lt;/em&gt; hear back from them, in the worst case scenario, you may discover
that a hard fork will be necessary because they don’t consider your issue
valid, but even that information will save you time, if you know it before you
get started.  In the best case, you may get a reply from the project telling
you that you’ve misunderstood its functionality and that there is already a
configuration parameter or usage pattern that will resolve your problems with
no new code.  But in all cases, you will benefit from early coordination on
&lt;em&gt;what&lt;/em&gt; needs fixing before you get to &lt;em&gt;how&lt;/em&gt; to fix it.&lt;/p&gt;
&lt;h3 id=step-1-source-code-and-ci-setup&gt;Step 1: Source Code and CI Setup&lt;/h3&gt;
&lt;p&gt;Fork the source code for your upstream dependency to a writable location where
it can live at least for the duration of this one bug-fix, and possibly for the
duration of your application’s use of the dependency.  After all, you might
want to fix more than &lt;em&gt;one&lt;/em&gt; bug in LibBar.&lt;/p&gt;
&lt;p&gt;You want to have a place where you can put your edits, that will be version
controlled and code reviewed according to your normal development process.
This probably means you’ll need to have your own main branch that diverges from
your upstream’s main branch.&lt;/p&gt;
&lt;p&gt;Remember: you’re going to need to deploy this to &lt;em&gt;your production&lt;/em&gt;, so testing
gates that your upstream only applies to final releases of LibBar will need to
be applied to every commit here.&lt;/p&gt;
&lt;p&gt;Depending on your LibBar’s own development process, this may result in slightly
unusual configurations where, for example, your fixes are written against the
last LibBar release tag, rather than its current&lt;sup id=fnref:5:dependency-cutout-workflow-pattern-2025-11&gt;&lt;a class=footnote-ref href=#fn:5:dependency-cutout-workflow-pattern-2025-11 id=fnref:5&gt;5&lt;/a&gt;&lt;/sup&gt; &lt;code&gt;main&lt;/code&gt;; if the project has a branch-freshness requirement, you
might need two branches, one for your upstream PR (based on main) and one for
your own use (based on the release branch with your changes).&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Ideally&lt;/em&gt; for projects with really good CI and a strong “keep main
release-ready at all times” policy, you can deploy straight from a development
branch, but it’s good to take a moment to consider this before you get started.
It’s usually easier to rebase changes from an older HEAD onto a newer one than
it is to go backwards.&lt;/p&gt;
&lt;p&gt;Speaking of CI, you will want to have your own CI system. The fact that GitHub
Actions has become a de-facto lingua franca of continuous integration means
that this step may be quite simple, and your forked repo can just run its own
instance.&lt;/p&gt;
&lt;h4 id=optional-bonus-step-1a-artifact-management&gt;Optional Bonus Step 1a: Artifact Management&lt;/h4&gt;
&lt;p&gt;If you have an in-house artifact repository, you should set that up for your
dependency too, and upload your own build artifacts to it.  You can often treat
your modified dependency as an extension of your own source tree and install
from a GitHub URL, but if you’ve already gone to the trouble of having an
in-house package repository, you can pretend you’ve taken over maintenance of
the upstream package temporarily (which you kind of have) and leverage those
workflows for caching and build-time savings as you would with any other
internal repo.&lt;/p&gt;
&lt;h3 id=step-2-do-the-fix&gt;Step 2: Do The Fix&lt;/h3&gt;
&lt;p&gt;Now that you’ve got somewhere to edit LibBar’s code, you will want to actually
fix the bug.&lt;/p&gt;
&lt;h4 id=step-2a-local-filesystem-setup&gt;Step 2a: Local Filesystem Setup&lt;/h4&gt;
&lt;p&gt;&lt;em&gt;Before&lt;/em&gt; you have a production version on your own deployed branch, you’ll want
to test locally, which means having &lt;em&gt;both&lt;/em&gt; repositories in a single integrated
development environment.&lt;/p&gt;
&lt;p&gt;At this point, you will want to have a &lt;em&gt;local filesystem reference&lt;/em&gt; to your
LibBar dependency, so that you can make real-time edits, without going through
a slow cycle of pushing to a branch in your LibBar fork, pushing to a FooApp
branch, and waiting for all of CI to run on both.&lt;/p&gt;
&lt;p&gt;This is useful in both directions: as you prepare the FooApp branch that makes
any necessary updates on that end, you’ll want to make sure that FooApp can
exercise the LibBar fix in any integration tests.  As you work on the LibBar
fix itself, you’ll also want to be able to use FooApp to exercise the code and
see if you’ve missed anything - and this, you wouldn’t get in CI, since LibBar
can’t depend on FooApp itself.&lt;/p&gt;
&lt;p&gt;In short, you want to be able to treat both projects as an integrated
&lt;em&gt;development environment&lt;/em&gt;, with support from your usual testing and debugging
tools, just as much as you want your deployment output to be an integrated
artifact.&lt;/p&gt;
&lt;h4 id=step-2b-branch-setup-for-pr&gt;Step 2b: Branch Setup for PR&lt;/h4&gt;
&lt;p&gt;However, for continuous integration to work, you will &lt;em&gt;also&lt;/em&gt; need to have a
remote resource reference of some kind from FooApp’s branch to LibBar.  You
will need 2 pull requests: the first to land your LibBar changes to your
internal LibBar fork and make sure it’s passing its &lt;em&gt;own&lt;/em&gt; tests, and then a
second PR to switch your LibBar dependency from the public repository to your
internal fork.&lt;/p&gt;
&lt;p&gt;At this step it is &lt;em&gt;very important&lt;/em&gt; to ensure that there is an issue filed on
your own internal backlog to drop your LibBar fork.  You do not want to lose
track of this work; it is technical debt that must be addressed.&lt;/p&gt;
&lt;p&gt;Until it’s addressed, automated tools like Dependabot will not be able to apply
security updates to LibBar for you; you’re going to need to manually integrate
every upstream change.  This type of work is itself very easy to drop or lose
track of, so you might just end up stuck on a vulnerable version.&lt;/p&gt;
&lt;h3 id=step-3-deploy-internally&gt;Step 3: Deploy Internally&lt;/h3&gt;
&lt;p&gt;Now that you’re confident that the fix will work, and that your
temporarily-internally-maintained version of LibBar isn’t going to break
anything on &lt;em&gt;your&lt;/em&gt; site, it’s time to deploy.&lt;/p&gt;
&lt;p&gt;Some &lt;a href="https://www.esa.int/Applications/Connectivity_and_Secure_Communications/Atlas_lifts_satcom_heritage#:~:text=They%20need%20proof%20that%20it%20has%20already%20worked%20in%20space%2C%20that%20it%20has%20‘flight%20heritage’"&gt;deployment
heritage&lt;/a&gt;
should help to provide &lt;em&gt;some&lt;/em&gt; evidence that your fix is ready to land in
LibBar, but at the next step, please remember that your production environment
isn’t necessarily emblematic of that of all LibBar users.&lt;/p&gt;
&lt;h3 id=step-4-propose-externally&gt;Step 4: Propose Externally&lt;/h3&gt;
&lt;p&gt;You’ve got the fix, you’ve tested the fix, you’ve got the fix in your own
production, you’ve told upstream you want to send them some changes.  Now, it’s
time to make the pull request.&lt;/p&gt;
&lt;p&gt;You’re likely going to get some feedback on the PR, even if you think it’s
already ready to go; as I said, despite having been proven in &lt;em&gt;your&lt;/em&gt; production
environment, you may get feedback about additional concerns from other users
that you’ll need to address before LibBar’s maintainers can land it.&lt;/p&gt;
&lt;p&gt;As you process the feedback, make sure that each new iteration of your branch
gets re-deployed to your own production. It would be a huge bummer to go
through all this trouble, and then end up unable to deploy the next publicly
released version of LibBar within FooApp because you forgot to test that your
responses to feedback &lt;em&gt;still worked&lt;/em&gt; on your own environment.&lt;/p&gt;
&lt;h4 id=step-4a-hurry-up-and-wait&gt;Step 4a: Hurry Up And Wait&lt;/h4&gt;
&lt;p&gt;If you’re lucky, upstream will land your changes to LibBar.  But, there’s still
no release version available.  Here, you’ll have to stay in a holding pattern
until upstream can finalize the release on their end.&lt;/p&gt;
&lt;p&gt;Depending on some particulars, it &lt;em&gt;might&lt;/em&gt; make sense at this point to archive
your internal LibBar repository and move your pinned release version to a git
hash of the LibBar version where your fix landed, in their repository.&lt;/p&gt;
&lt;p&gt;Before you do this, check in with the LibBar core team and make sure that they
understand that’s what you’re doing and they don’t have any wacky workflows
which may involve rebasing or eliding that commit as part of their release
process.&lt;/p&gt;
&lt;h3 id=step-5-unwind-everything&gt;Step 5: Unwind Everything&lt;/h3&gt;
&lt;p&gt;Finally, you eventually want to stop carrying any patches and move back to an
official released version that integrates your fix.&lt;/p&gt;
&lt;p&gt;You want to do this because this is what the upstream will expect when you are
reporting bugs.  Part of the benefit of using open source is benefiting from
the collective work to do bug-fixes and such, so you don’t want to be stuck off
on a pinned git hash that the developers do not support for anyone else.&lt;/p&gt;
&lt;p&gt;As I said in step 2b&lt;sup id=fnref:6:dependency-cutout-workflow-pattern-2025-11&gt;&lt;a class=footnote-ref href=#fn:6:dependency-cutout-workflow-pattern-2025-11 id=fnref:6&gt;6&lt;/a&gt;&lt;/sup&gt;, make sure to &lt;em&gt;maintain a tracking task&lt;/em&gt; for doing this
work, because leaving this sort of relatively &lt;em&gt;easy&lt;/em&gt;-to-clean-up technical debt
lying around is something that can potentially create a lot of aggravation for
no particular benefit.  Make sure to put your internal LibBar repository into
an appropriate state at this point as well.&lt;/p&gt;
&lt;h2 id=up-next&gt;Up Next&lt;/h2&gt;
&lt;p&gt;This is part 1 of a 2-part series.  In part 2, I will explore in depth how to
execute this workflow specifically for Python packages, using some popular
tools.  I’ll discuss my own workflow, standards like PEP 517 and
&lt;code&gt;pyproject.toml&lt;/code&gt;, and of course, by the popular demand that I just &lt;em&gt;know&lt;/em&gt; will
come, &lt;a href="https://github.com/astral-sh/uv"&gt;&lt;code&gt;uv&lt;/code&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=acknowledgments&gt;Acknowledgments&lt;/h2&gt;
&lt;p class=update-note&gt;Thank you to &lt;a href="/pages/patrons.html"&gt;my patrons&lt;/a&gt; who are supporting my writing on
this blog.  If you like what you’ve read here and you’d like to read more of
it, or you’d like to support my &lt;a href="https://github.com/glyph/"&gt;various open-source
endeavors&lt;/a&gt;, you can &lt;a href="/pages/patrons.html"&gt;support my work as a
sponsor&lt;/a&gt;!&lt;/p&gt;
&lt;div class=footnote&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=fn:1:dependency-cutout-workflow-pattern-2025-11&gt;
&lt;p id=fn:1&gt;if you already have all the tooling associated with a monorepo,
&lt;em&gt;including&lt;/em&gt; the ability to manage divergence and reintegrate patches with
upstream, you already have the higher-overhead version of the workflow I am
going to propose, so, never mind. but chances are you don’t have that, very
few companies do. &lt;a class=footnote-backref href=#fnref:1:dependency-cutout-workflow-pattern-2025-11 title="Jump back to footnote 1 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:2:dependency-cutout-workflow-pattern-2025-11&gt;
&lt;p id=fn:2&gt;In any business where one must wrangle with Legal, 3 hours is a &lt;em&gt;wildly&lt;/em&gt;
optimistic estimate. &lt;a class=footnote-backref href=#fnref:2:dependency-cutout-workflow-pattern-2025-11 title="Jump back to footnote 2 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:3:dependency-cutout-workflow-pattern-2025-11&gt;
&lt;p id=fn:3&gt;&lt;a href="https://mastodon.social/@mcc/112117339397138167"&gt;c.f. &lt;code&gt;@mcc@mastodon.social&lt;/code&gt;&lt;/a&gt; &lt;a class=footnote-backref href=#fnref:3:dependency-cutout-workflow-pattern-2025-11 title="Jump back to footnote 3 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:4:dependency-cutout-workflow-pattern-2025-11&gt;
&lt;p id=fn:4&gt;&lt;a href="https://mastodon.social/@geofft/112186487032016599"&gt;c.f. &lt;code&gt;@geofft@mastodon.social&lt;/code&gt;&lt;/a&gt; &lt;a class=footnote-backref href=#fnref:4:dependency-cutout-workflow-pattern-2025-11 title="Jump back to footnote 4 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:5:dependency-cutout-workflow-pattern-2025-11&gt;
&lt;p id=fn:5&gt;In an ideal world every project would &lt;a href="https://martinfowler.com/articles/continuousIntegration.html#FixBrokenBuildsImmediately"&gt;keep its main branch ready to
release at all times, no matter
what&lt;/a&gt;
but we do not live in an ideal world. &lt;a class=footnote-backref href=#fnref:5:dependency-cutout-workflow-pattern-2025-11 title="Jump back to footnote 5 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:6:dependency-cutout-workflow-pattern-2025-11&gt;
&lt;p id=fn:6&gt;In this case, there is no question. It’s 2b only, no not-2b. &lt;a class=footnote-backref href=#fnref:6:dependency-cutout-workflow-pattern-2025-11 title="Jump back to footnote 6 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;&lt;/body&gt;</content><category term="misc"></category><category term="programming"></category><category term="python"></category><category term="deployment"></category><category term="open-source"></category></entry><entry><title>The Futzing Fraction</title><link href="https://blog.glyph.im/2025/08/futzing-fraction.html" rel="alternate"></link><published>2025-08-15T00:51:00-07:00</published><updated>2025-08-15T00:51:00-07:00</updated><author><name>Glyph</name></author><id>tag:blog.glyph.im,2025-08-15:/2025/08/futzing-fraction.html</id><summary type="html">&lt;p&gt;At least &lt;em&gt;some&lt;/em&gt; of your time with genAI will be spent just kind of… futzing with it.&lt;/p&gt;</summary><content type="html">&lt;body&gt;&lt;p&gt;The most optimistic vision of generative AI&lt;sup id=fnref:1:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:1:futzing-fraction-2025-8 id=fnref:1&gt;1&lt;/a&gt;&lt;/sup&gt; is that it will relieve us of
the tedious, repetitive elements of knowledge work so that we can get to work
on the &lt;em&gt;really&lt;/em&gt; interesting problems that such tedium stands in the way of.
Even if you fully believe in this vision, it’s hard to deny that &lt;em&gt;today&lt;/em&gt;, some
tedium is associated with the process of using generative AI itself.&lt;/p&gt;
&lt;p&gt;Generative AI also
&lt;a href="https://futurism.com/the-byte/openai-chatgpt-pro-subscription-losing-money"&gt;isn’t&lt;/a&gt;
&lt;a href="https://www.computerworld.com/article/4021954/rushing-into-genai-prepare-for-budget-blowouts-and-broken-promises.html"&gt;free&lt;/a&gt;,
and so, as responsible consumers, we need to ask: is it worth it?  What’s the
&lt;a href="https://www.investopedia.com/articles/basics/10/guide-to-calculating-roi.asp"&gt;ROI&lt;/a&gt;
of genAI, and how can we tell?  In this post, I’d like to explore a logical
framework for evaluating genAI expenditures, to determine if your organization
is getting its money’s worth.&lt;/p&gt;
&lt;h1 id=perpetually-proffering-permuted-prompts&gt;Perpetually Proffering Permuted Prompts&lt;/h1&gt;
&lt;p&gt;I think most LLM users would agree with me that a typical workflow with an LLM
rarely involves prompting it only one time and getting a perfectly useful
answer that solves the whole problem.&lt;/p&gt;
&lt;p&gt;Generative AI best practices, even &lt;a href="https://techcommunity.microsoft.com/blog/azuredevcommunityblog/evaluating-generative-ai-best-practices-for-developers/4271488#:~:text=Frequent%20and%20scheduled%20evaluations%20should%20be%20embedded%20into%20the%20development%20cycle"&gt;from the most optimistic
vendors&lt;/a&gt;
all suggest that you should continuously evaluate everything.  ChatGPT, which
is really the
&lt;a href="https://www.wheresyoured.at/the-haters-gui/#:~:text=ChatGPT%20has%20500%20million%20weekly%20users%2C%20and%20otherwise%2C%20it%20seems%20that%20other%20services%20struggle%20to%20get%2015%20million%20of%20them"&gt;only&lt;/a&gt;
genAI product with significantly scaled adoption, still says at the bottom of
&lt;em&gt;every&lt;/em&gt; interaction:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;&lt;em&gt;ChatGPT can make mistakes. Check important info.&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;If we have to “check important info” on every interaction, it stands to reason
that even if we think it’s useful, &lt;em&gt;some&lt;/em&gt; of those checks will find an error.
Again, if we think it’s useful, presumably the next thing to do is to perturb
our prompt somehow, and issue it again, in the hopes that the next invocation
will, by dint of either:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href="https://pivot-to-ai.com/2025/06/05/generative-ai-runs-on-gambling-addiction-just-one-more-prompt-bro/"&gt;better luck this
time&lt;/a&gt;
   with the &lt;a href="https://en.wikipedia.org/wiki/Stochastic_parrot"&gt;stochastic&lt;/a&gt; aspect of the inference process,&lt;/li&gt;
&lt;li&gt;enhanced application of our skill to
   &lt;a href="https://www.fastcompany.com/91327911/prompt-engineering-going-extinct"&gt;engineer&lt;/a&gt;
   a better prompt based on the deficiencies of the current inference, or&lt;/li&gt;
&lt;li&gt;better performance of the model by populating additional
   &lt;a href="https://research.trychroma.com/context-rot"&gt;context&lt;/a&gt; in subsequent chained
   prompts.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Unfortunately, given the relative lack of &lt;em&gt;reliable&lt;/em&gt; methods to re-generate the
prompt and receive a better answer&lt;sup id=fnref:2:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:2:futzing-fraction-2025-8 id=fnref:2&gt;2&lt;/a&gt;&lt;/sup&gt;, checking the output and re-prompting
the model can feel like just kinda futzing around with it.  You try, you get a
wrong answer, you try a few more times, eventually you get the right answer
that you wanted in the first place.  It’s a somewhat unsatisfying process, but
if you get the right answer eventually, it does feel like progress, and you
didn’t need to use up another human’s time.&lt;/p&gt;
&lt;p&gt;In fact, the hottest buzzword of the last hype cycle is “agentic”.  While I
have my own feelings about this particular word&lt;sup id=fnref:3:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:3:futzing-fraction-2025-8 id=fnref:3&gt;3&lt;/a&gt;&lt;/sup&gt;, its current &lt;em&gt;practical&lt;/em&gt;
definition is “a generative AI system which automates the process of
re-prompting itself, by having a deterministic program evaluate its outputs for
correctness”.&lt;/p&gt;
&lt;p&gt;A better term for an “agentic” system would be a “self-futzing system”.&lt;/p&gt;
&lt;p&gt;However, the ability to automate &lt;em&gt;some&lt;/em&gt; level of checking and re-prompting does
not mean that you can &lt;em&gt;fully&lt;/em&gt; delegate tasks to an agentic tool, either.  It
is, plainly put, not safe. If you leave the AI on its own, you will get
&lt;em&gt;terrible&lt;/em&gt; results that will at best make for a funny story&lt;sup id=fnref:4:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:4:futzing-fraction-2025-8 id=fnref:4&gt;4&lt;/a&gt;&lt;/sup&gt;&lt;sup id=fnref:5:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:5:futzing-fraction-2025-8 id=fnref:5&gt;5&lt;/a&gt;&lt;/sup&gt; and at
worst might end up causing serious damage&lt;sup id=fnref:6:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:6:futzing-fraction-2025-8 id=fnref:6&gt;6&lt;/a&gt;&lt;/sup&gt;&lt;sup id=fnref:7:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:7:futzing-fraction-2025-8 id=fnref:7&gt;7&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;Taken together, this all means that for &lt;em&gt;any&lt;/em&gt; consequential task that you want
to accomplish with genAI, you need an expert &lt;a href="https://en.wikipedia.org/wiki/Human-in-the-loop"&gt;human in the
loop&lt;/a&gt;.  The human must be
capable of independently doing the job that the genAI system is being asked to
accomplish.&lt;/p&gt;
&lt;p&gt;When the genAI guesses correctly and produces usable output, some of the
human’s time will be saved.  When the genAI guesses wrong and produces
hallucinatory gibberish or even “correct” output that nevertheless fails to
account for some unstated but necessary property such as security or scale,
some of the human’s time will be wasted evaluating it and re-trying it.&lt;/p&gt;
&lt;h1 id=income-from-investment-in-inference&gt;Income from Investment in Inference&lt;/h1&gt;
&lt;p&gt;Let’s evaluate an abstract, hypothetical genAI system that can automate some
work for our organization.  To avoid implicating any specific vendor, let’s
call the system “Mallory”.&lt;/p&gt;
&lt;p&gt;Is Mallory worth the money?  How can we know?&lt;/p&gt;
&lt;p&gt;Logically, there are only two outcomes that might result from using Mallory to
do our work.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;We prompt Mallory to do some work; we check its work, it is correct, and
   some time is saved.&lt;/li&gt;
&lt;li&gt;We prompt Mallory to do some work; we check its work, it fails, and we futz
   around with the result; this time is wasted.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;As a &lt;em&gt;logical&lt;/em&gt; framework, this makes sense, but ROI is an arithmetical concept,
not a logical one.  So let’s translate this into some terms.&lt;/p&gt;
&lt;p&gt;In order to evaluate Mallory, let’s define the Futzing Fraction, “&lt;math&gt;
&lt;mi&gt;FF&lt;/mi&gt; &lt;/math&gt;”, in terms of the following variables:&lt;/p&gt;
&lt;dl&gt;
&lt;dt&gt;&lt;math&gt;&lt;mi&gt;H&lt;/mi&gt;&lt;/math&gt;&lt;/dt&gt;
&lt;dd&gt;
&lt;p&gt;the average amount of time a &lt;strong&gt;&lt;em&gt;H&lt;/em&gt;&lt;/strong&gt;uman worker would take to do a task,
unaided by Mallory&lt;/p&gt;
&lt;/dd&gt;
&lt;dt&gt;&lt;math&gt;&lt;mi&gt;I&lt;/mi&gt;&lt;/math&gt;&lt;/dt&gt;
&lt;dd&gt;
&lt;p&gt;the amount of time that Mallory takes to run one &lt;strong&gt;&lt;em&gt;I&lt;/em&gt;&lt;/strong&gt;nference&lt;sup id=fnref:8:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:8:futzing-fraction-2025-8 id=fnref:8&gt;8&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/dd&gt;
&lt;dt&gt;&lt;math&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;/math&gt;&lt;/dt&gt;
&lt;dd&gt;
&lt;p&gt;the amount of time that a human has to spend &lt;strong&gt;&lt;em&gt;C&lt;/em&gt;&lt;/strong&gt;hecking Mallory’s output for
each inference&lt;/p&gt;
&lt;/dd&gt;
&lt;dt&gt;&lt;math&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;/math&gt;&lt;/dt&gt;
&lt;dd&gt;
&lt;p&gt;the &lt;strong&gt;&lt;em&gt;P&lt;/em&gt;&lt;/strong&gt;robability that Mallory will produce a correct inference for each prompt&lt;/p&gt;
&lt;/dd&gt;
&lt;dt&gt;&lt;math&gt;&lt;mi&gt;W&lt;/mi&gt;&lt;/math&gt;&lt;/dt&gt;
&lt;dd&gt;
&lt;p&gt;the average amount of time that it takes for a human to &lt;strong&gt;&lt;em&gt;W&lt;/em&gt;&lt;/strong&gt;rite one prompt for
Mallory&lt;/p&gt;
&lt;/dd&gt;
&lt;dt&gt;&lt;math&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/math&gt;&lt;/dt&gt;
&lt;dd&gt;
&lt;p&gt;since we are normalizing everything to &lt;em&gt;time&lt;/em&gt;, rather than &lt;em&gt;money&lt;/em&gt;, we do also have to account for the dollar of Mallory as as a product, so we will include the &lt;strong&gt;&lt;em&gt;E&lt;/em&gt;&lt;/strong&gt;quivalent amount of human time we could purchase for the marginal cost of one&lt;sup id=fnref:9:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:9:futzing-fraction-2025-8 id=fnref:9&gt;9&lt;/a&gt;&lt;/sup&gt; inference.&lt;/p&gt;
&lt;/dd&gt;
&lt;/dl&gt;
&lt;p&gt;As in last week’s example of &lt;a href="https://blog.glyph.im/2025/08/r0mls-ratio.html"&gt;simple ROI
arithmetic&lt;/a&gt;, we will put our costs in the
numerator, and our benefits in the denominator.&lt;/p&gt;
&lt;div style="font-size: 30px; text-align: center;"&gt;
&lt;math&gt;
    &lt;mi&gt;FF&lt;/mi&gt; &lt;mo&gt; = &lt;/mo&gt;
    &lt;mfrac&gt;
        &lt;mrow&gt;&lt;mi&gt;W&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;I&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;
        &lt;mrow&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;!--&lt;mo&gt;✕&lt;/mo&gt;--&gt; &lt;mi&gt;H&lt;/mi&gt;&lt;/mrow&gt;
    &lt;/mfrac&gt;
&lt;/math&gt;
&lt;/div&gt;

&lt;p&gt;The idea here is that for each prompt, the &lt;em&gt;minimum&lt;/em&gt; amount of time-equivalent cost possible is &lt;math&gt;&lt;mi&gt;W&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;I&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/math&gt;.  The user must, at least once, write a prompt, wait for inference to run, then check the output; and, of course, pay any costs to Mallory’s vendor.&lt;/p&gt;
&lt;p&gt;If the probability of a correct answer is &lt;math&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mn&gt;3&lt;/mn&gt;&lt;/mfrac&gt;&lt;/math&gt;, then they will do this entire process 3 times&lt;sup id=fnref:10:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:10:futzing-fraction-2025-8 id=fnref:10&gt;10&lt;/a&gt;&lt;/sup&gt;, so we put &lt;math&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;/math&gt; in the denominator.  Finally, we divide everything by &lt;math&gt;&lt;mi&gt;H&lt;/mi&gt;&lt;/math&gt;, because we are trying to determine if we are actually saving any time or money, versus just letting our existing human, who has to be driving this process anyway, do the whole thing.&lt;/p&gt;
&lt;p&gt;If the Futzing Fraction evaluates to a number greater than 1, &lt;a href="https://blog.glyph.im/2025/08/r0mls-ratio.html"&gt;as previously discussed, you are a bozo&lt;/a&gt;; you’re spending more time futzing with Mallory than getting value out of it.&lt;/p&gt;
&lt;h1 id=figuring-out-the-fraction-is-frustrating&gt;Figuring out the Fraction is Frustrating&lt;/h1&gt;
&lt;p&gt;In order to even evaluate the value of the Futzing Fraction though, you have to
have a sound method to even get a vague sense of all the terms.&lt;/p&gt;
&lt;p&gt;If you are a business leader, a lot of this is relatively easy to measure.  You
vaguely know what &lt;math&gt;&lt;mi&gt;H&lt;/mi&gt;&lt;/math&gt; is, because you know what your
payroll costs, and similarly, you can figure out &lt;math&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/math&gt; with
some pretty trivial arithmetic based on Mallory’s pricing table.  There are endless
YouTube channels, spec sheets and benchmarks to give you &lt;math&gt;&lt;mi&gt;I&lt;/mi&gt;&lt;/math&gt;. &lt;math&gt;&lt;mi&gt;W&lt;/mi&gt;&lt;/math&gt; is probably going to be so small compared to &lt;math&gt;&lt;mi&gt;H&lt;/mi&gt;&lt;/math&gt; that it hardly merits consideration&lt;sup id=fnref:11:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:11:futzing-fraction-2025-8 id=fnref:11&gt;11&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;But, are you measuring &lt;math&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;/math&gt;?  If your employees are &lt;em&gt;not&lt;/em&gt; checking the outputs of the AI, you’re on a path to catastrophe that no ROI calculation can capture, so it had &lt;em&gt;better&lt;/em&gt; be greater than zero.&lt;/p&gt;
&lt;p&gt;Are you measuring &lt;math&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;/math&gt;?  How often does the AI get it right on the first try?&lt;/p&gt;
&lt;h2 id=challenges-to-computing-checking-costs&gt;Challenges to Computing Checking Costs&lt;/h2&gt;
&lt;p&gt;In the fraction defined above, the term &lt;math&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;/math&gt; is going to be
large.  Larger than you think.&lt;/p&gt;
&lt;p&gt;Measuring &lt;math&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;/math&gt; and &lt;math&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;/math&gt; with a high
degree of precision is probably going to be very hard; possibly unreasonably
so, or too expensive&lt;sup id=fnref:12:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:12:futzing-fraction-2025-8 id=fnref:12&gt;12&lt;/a&gt;&lt;/sup&gt; to bother with in practice.  So you will undoubtedly need
to work with estimates and proxy metrics.  But you have to be aware that this
is a problem domain where your normal method of estimating is going to be
&lt;em&gt;extremely&lt;/em&gt; vulnerable to inherent cognitive bias, and find ways to measure.&lt;/p&gt;
&lt;h3 id=margins-money-and-metacognition&gt;Margins, Money, and Metacognition&lt;/h3&gt;
&lt;p&gt;First let’s discuss cognitive and metacognitive bias.&lt;/p&gt;
&lt;p&gt;My favorite cognitive bias is the &lt;a href="https://en.wikipedia.org/wiki/Availability_heuristic"&gt;availability
heuristic&lt;/a&gt; and a close
second is its cousin &lt;a href="https://en.wikipedia.org/wiki/Salience_(neuroscience)#Salience_bias"&gt;salience
bias&lt;/a&gt;.
Humans are empirically predisposed towards noticing and remembering things that
are more striking, and to overestimate their frequency.&lt;/p&gt;
&lt;p&gt;If you are estimating the variables above based on the &lt;em&gt;vibe&lt;/em&gt; that you’re
getting from the experience of using an LLM, you may be overestimating its
utility.&lt;/p&gt;
&lt;p&gt;Consider a slot machine.&lt;/p&gt;
&lt;p&gt;If you put a dollar in to a slot machine, and you lose that dollar, this is an
unremarkable event. Expected, even.  It doesn’t seem interesting.  You can
repeat this over and over again, a thousand times, and each time it will seem
equally unremarkable.  If you do it a thousand times, you will probably get
gradually more anxious as your sense of your dwindling bank account becomes
slowly more salient, but losing one more dollar still seems unremarkable.&lt;/p&gt;
&lt;p&gt;If you put a dollar in a slot machine and it gives you a &lt;em&gt;thousand&lt;/em&gt; dollars,
that will probably seem pretty cool.  Interesting.  Memorable.  You might tell
a story about this happening, but you definitely wouldn’t really remember any
particular time you lost one dollar.&lt;/p&gt;
&lt;p&gt;Luckily, when you arrive at a casino with slot machines, you probably know well
enough to set a hard budget in the form of some amount of physical currency you
will have available to you.  The odds are against you, you’ll probably lose it
all, but any responsible gambler will have an immediate, physical
representation of their balance in front of them, so when they have lost it
all, they can see that their hands are empty, and can try to resist the “just
one more pull” temptation, after hitting that limit.&lt;/p&gt;
&lt;p&gt;Now, consider Mallory.&lt;/p&gt;
&lt;p&gt;If you put ten minutes into writing a prompt, and Mallory gives a completely
off-the-rails, useless answer, and you lose ten minutes, well, that’s just what
using a computer is like sometimes.  Mallory malfunctioned, or hallucinated,
but it does that sometimes, everybody knows that.  You only wasted ten minutes.
It’s fine.  Not a big deal.  Let’s try it a few more times.  Just ten more
minutes.  It’ll probably work this time.&lt;/p&gt;
&lt;p&gt;If you put ten minutes into writing a prompt, and it completes a task that
would have otherwise taken you 4 hours, that feels amazing.  Like the computer
is &lt;em&gt;magic&lt;/em&gt;! An absolute endorphin rush.&lt;/p&gt;
&lt;p&gt;Very memorable.  When it happens, it feels like &lt;math&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/math&gt;.&lt;/p&gt;
&lt;p&gt;But... did you have a time budget before you started?  Did you have a specified
N such that “I will give up on Mallory as soon as I have spent N minutes
attempting to solve this problem with it”?  When the jackpot finally pays out
that 4 hours, did you &lt;em&gt;notice&lt;/em&gt; that you put 6 hours worth of 10-minute prompt
coins into it in?&lt;/p&gt;
&lt;p&gt;If you are attempting to use the same sort of heuristic intuition that probably
works &lt;em&gt;pretty well&lt;/em&gt; for other business leadership decisions, Mallory’s
slot-machine chat-prompt user interface is practically &lt;em&gt;designed&lt;/em&gt; to subvert
those sensibilities.  Most business activities do not have nearly such an
emotionally variable, intermittent reward schedule.  They’re not going to trick
you with this sort of cognitive illusion.&lt;/p&gt;
&lt;p&gt;Thus far we have been talking about cognitive bias, but there is a
metacognitive bias at play too: while
&lt;a href="https://en.wikipedia.org/wiki/Dunning–Kruger_effect"&gt;Dunning-Kruger&lt;/a&gt;,
everybody’s favorite metacognitive bias does have some
&lt;a href="https://www.sciencedirect.com/science/article/pii/S1877042814051489#bib0040"&gt;problems&lt;/a&gt;
with it, the main underlying metacognitive bias is that we tend to &lt;em&gt;believe our
own thoughts and perceptions&lt;/em&gt;, and it requires active effort to distance
ourselves from them, even if we know they might be wrong.&lt;/p&gt;
&lt;p&gt;This means you must assume any &lt;em&gt;intuitive&lt;/em&gt; estimate of &lt;math&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;/math&gt;
is going to be biased low; similarly &lt;math&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;/math&gt; is going to be
biased high.  You will forget the time you spent checking, and you will
underestimate the number of times you had to re-check.&lt;/p&gt;
&lt;p&gt;To avoid this, you will need to decide on a &lt;a href="https://en.wikipedia.org/wiki/Ulysses_pact"&gt;Ulysses
pact&lt;/a&gt; to provide some inputs to a
calculation for these factors that you will not be able to able to fudge if
they seem wrong to you.&lt;/p&gt;
&lt;h3 id=problematically-plausible-presentation&gt;Problematically Plausible Presentation&lt;/h3&gt;
&lt;p&gt;Another nasty little cognitive-bias landmine for you to watch out for is the
&lt;a href="https://en.wikipedia.org/wiki/Authority_bias"&gt;authority bias&lt;/a&gt;, for two
reasons:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;People will tend to see Mallory as an unbiased, external authority, and
   thereby see it as more of an authority than a similarly-situated human&lt;sup id=fnref:13:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:13:futzing-fraction-2025-8 id=fnref:13&gt;13&lt;/a&gt;&lt;/sup&gt;.&lt;/li&gt;
&lt;li&gt;Being an LLM, Mallory will be overconfident in its answers&lt;sup id=fnref:14:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:14:futzing-fraction-2025-8 id=fnref:14&gt;14&lt;/a&gt;&lt;/sup&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The nature of LLM training is also such that commonly co-occurring tokens in
the training corpus produce higher likelihood of co-occurring in the output;
they’re just going to be closer together in the vector-space of the weights;
that’s, like, what training a model &lt;em&gt;is&lt;/em&gt;, establishing those relationships.&lt;/p&gt;
&lt;p&gt;If you’ve ever used an heuristic to informally evaluate someone’s credibility
by listening for industry-specific shibboleths or ways of describing a
particular issue, that skill is now useless.  Having ingested every industry’s
expert literature, commonly-occurring phrases will always be present in
Mallory’s output.  Mallory will usually sound like an expert, but then make
mistakes at random.&lt;sup id=fnref:15:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:15:futzing-fraction-2025-8 id=fnref:15&gt;15&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;While you might intuitively estimate &lt;math&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;/math&gt; by thinking “well,
if I asked a &lt;em&gt;person&lt;/em&gt;, how could I check that &lt;em&gt;they&lt;/em&gt; were correct, and how long
would that take?” that estimate will be extremely optimistic, because the
heuristic techniques you would use to quickly evaluate incorrect information
from other humans will fail with Mallory.  You need to go all the way back to
primary sources and actually &lt;em&gt;fully&lt;/em&gt; verify the output every time, or you will
likely fall into one of these traps.&lt;/p&gt;
&lt;h3 id=mallory-mangling-mentorship&gt;Mallory Mangling Mentorship&lt;/h3&gt;
&lt;p&gt;So far, I’ve been describing the effect Mallory will have in the context of an
individual attempting to get some work done. If we are considering
organization-wide adoption of Mallory, however, we must &lt;em&gt;also&lt;/em&gt; consider the
impact on team dynamics.  There are a number of possible potential side effects
that one might consider when looking at, but here I will focus on just one that
I have observed.&lt;/p&gt;
&lt;p&gt;I have a cohort of friends in the software industry, most of whom are
individual contributors.  I’m a programmer who likes programming, so are most
of my friends, and we are also (&lt;strong&gt;sigh&lt;/strong&gt;), charitably, &lt;em&gt;pretty solidly
middle-aged&lt;/em&gt; at this point, so we tend to have a lot of experience.&lt;/p&gt;
&lt;p&gt;As such, we are often the folks that the team — or, in my case, the community —
goes to when less-experienced folks need answers.&lt;/p&gt;
&lt;p&gt;On its own, this is actually pretty great.  Answering questions from more
junior folks is one of the best parts of a software development job.  It’s an
opportunity to be helpful, mostly just by knowing a thing we already knew.  And
it’s an opportunity to help someone else improve their own agency by giving
them knowledge that they can use in the future.&lt;/p&gt;
&lt;p&gt;However, generative AI throws a bit of a wrench into the mix.&lt;/p&gt;
&lt;p&gt;Let’s imagine a scenario where we have 2 developers: Alice, a staff engineer
who has a good understanding of the system being built, and Bob, a relatively
junior engineer who is still onboarding.&lt;/p&gt;
&lt;p&gt;The traditional interaction between Alice and Bob, when Bob has a question,
goes like this:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Bob gets confused about something in the system being developed, because
   Bob’s understanding of the system is incorrect.&lt;/li&gt;
&lt;li&gt;Bob formulates a question based on this confusion.&lt;/li&gt;
&lt;li&gt;Bob asks Alice that question.&lt;/li&gt;
&lt;li&gt;Alice knows the system, so she gives an answer which
   accurately reflects the state of the system to Bob.&lt;/li&gt;
&lt;li&gt;Bob’s understanding of the system improves, and thus he will have fewer and
   better-informed questions going forward.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;You can imagine how repeating this simple 5-step process will eventually
transform Bob into a senior developer, and then he can start answering
questions on his own.  Making sufficient time for regularly iterating this loop
is the heart of any good mentorship process.&lt;/p&gt;
&lt;p&gt;Now, though, with Mallory in the mix, the process now has a new decision point,
changing it from a linear sequence to a flow chart.&lt;/p&gt;
&lt;p&gt;We begin the same way, with steps 1 and 2.  Bob’s confused, Bob formulates a
question, but then:&lt;/p&gt;
&lt;ol start=3&gt;
&lt;li&gt;Bob asks Mallory that question.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Here, our path then diverges into a “happy” path, a “meh” path, and a “sad”
path.&lt;/p&gt;
&lt;p&gt;The “happy” path proceeds like so:&lt;/p&gt;
&lt;ol start=4&gt;
&lt;li&gt;Mallory happens to formulate a correct answer.&lt;/li&gt;
&lt;li&gt;Bob’s understanding of the system improves, and thus he will have fewer and
   better-informed questions going forward.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Great. Problem solved. We just saved some of Alice’s time. But as we learned earlier,&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;em&gt;Mallory can make mistakes&lt;/em&gt;&lt;/strong&gt;.  When that happens, we will need to &lt;strong&gt;&lt;em&gt;check
important info&lt;/em&gt;&lt;/strong&gt;.  So let’s get checking:&lt;/p&gt;
&lt;ol start=4&gt;
&lt;li&gt;Mallory happens to formulate an &lt;em&gt;incorrect&lt;/em&gt; answer.&lt;/li&gt;
&lt;li&gt;Bob investigates this answer.&lt;/li&gt;
&lt;li&gt;Bob realizes that this answer is incorrect because it is inconsistent with
   some of his prior, correct knowledge of the system, or his investigation.&lt;/li&gt;
&lt;li&gt;Bob asks Alice the same question; GOTO traditional interaction step 4.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;On this path, Bob spent a while futzing around with Mallory, to no particular
benefit.  This wastes some of Bob’s time, but then again, Bob &lt;em&gt;could&lt;/em&gt; have
ended up on the happy path, so perhaps it was worth the risk; at least Bob
wasn’t wasting any of &lt;em&gt;Alice’s&lt;/em&gt; much more valuable time in the process.&lt;sup id=fnref:16:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:16:futzing-fraction-2025-8 id=fnref:16&gt;16&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;Notice that beginning at the start of step 4, we must begin allocating all of
Bob’s time to &lt;math&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;/math&gt;, so &lt;math&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;/math&gt; already
starts getting a bit bigger than if it were just Bob checking Mallory’s output
specifically on &lt;em&gt;tasks&lt;/em&gt; that Bob is doing.&lt;/p&gt;
&lt;p&gt;That brings us to the “sad” path.&lt;/p&gt;
&lt;ol start=4&gt;
&lt;li&gt;Mallory happens to formulate an &lt;em&gt;incorrect&lt;/em&gt; answer.&lt;/li&gt;
&lt;li&gt;Bob investigates this answer.&lt;/li&gt;
&lt;li&gt;Bob &lt;em&gt;does not realize&lt;/em&gt; that this answer is incorrect because he is unable to
   recognize any inconsistencies with his existing, incomplete knowledge of the
   system.&lt;/li&gt;
&lt;li&gt;Bob integrates Mallory’s incorrect information of the system into his mental
   model.&lt;/li&gt;
&lt;li&gt;Bob proceeds to make a larger and larger mess of his work, based on an
   incorrect mental model.&lt;/li&gt;
&lt;li&gt;Eventually, Bob asks Alice a new, worse question, based on this incorrect
   understanding.&lt;/li&gt;
&lt;li&gt;Sadly we &lt;em&gt;cannot&lt;/em&gt; return to the happy path at this point, because now Alice
    must unravel the complex series of confusing misunderstandings that Mallory
    has unfortunately conveyed to Bob at this point.  In the &lt;em&gt;really&lt;/em&gt; sad
    case, Bob actually &lt;em&gt;doesn’t believe&lt;/em&gt; Alice for a while, because Mallory
    seems unbiased&lt;sup id=fnref:17:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:17:futzing-fraction-2025-8 id=fnref:17&gt;17&lt;/a&gt;&lt;/sup&gt;, and Alice has to waste even more time &lt;em&gt;convincing&lt;/em&gt; Bob
    before she can simply &lt;em&gt;explain&lt;/em&gt; to him.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Now, we have wasted some of Bob’s time, &lt;em&gt;and&lt;/em&gt; some of Alice’s time.  Everything
from step 5-10 is &lt;math&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;/math&gt;, and as soon as Alice gets involved,
we are now adding to &lt;math&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;/math&gt; at &lt;em&gt;double&lt;/em&gt; real-time.  If more
team members are pulled in to the investigation, you are now multiplying &lt;math&gt;
&lt;mi&gt;C&lt;/mi&gt;&lt;/math&gt; by the number of investigators, potentially running at triple
or quadruple real time.&lt;/p&gt;
&lt;h3 id=but-thats-not-all&gt;But That’s Not All&lt;/h3&gt;
&lt;p&gt;Here I’ve presented a &lt;em&gt;brief&lt;/em&gt; selection reasons why &lt;math&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;/math&gt;
will be both large, and larger than you expect. To review:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Gambling-style mechanics of the user interface will interfere with your own
   self-monitoring and developing a good estimate.&lt;/li&gt;
&lt;li&gt;You can’t use human heuristics for quickly spotting bad answers.&lt;/li&gt;
&lt;li&gt;Wrong answers given to junior people who can’t evaluate them will waste more
   time from your more senior employees.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;But this is a &lt;em&gt;small&lt;/em&gt; selection of ways that Mallory’s output can cost you
money and time.  It’s harder to simplistically model second-order effects like
this, but there’s also a broad range of possibilities for ways that, rather
than simply checking and catching errors, an error slips through and starts
doing damage. Or ways in which the output isn’t exactly &lt;em&gt;wrong&lt;/em&gt;, but still
sub-optimal in ways which can be difficult to notice in the short term.&lt;/p&gt;
&lt;p&gt;For example, you might successfully vibe-code your way to launch a series of
applications, successfully “checking” the output along the way, but then
discover that the resulting code is unmaintainable garbage that prevents future
feature delivery, and needs to be re-written&lt;sup id=fnref:18:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:18:futzing-fraction-2025-8 id=fnref:18&gt;18&lt;/a&gt;&lt;/sup&gt;.  But this kind of
intellectual debt isn’t even specific to technical debt while coding; it can
even affect such apparently genAI-amenable fields as LinkedIn content
marketing&lt;sup id=fnref:19:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:19:futzing-fraction-2025-8 id=fnref:19&gt;19&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;h2 id=problems-with-the-prediction-of-p&gt;Problems with the Prediction of &lt;math&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;/math&gt;&lt;/h2&gt;
&lt;p&gt;&lt;math&gt; &lt;mi&gt;C&lt;/mi&gt; &lt;/math&gt; isn’t the only challenging term
though. &lt;math&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;/math&gt;, is just as, if not more important, and just as
hard to measure.&lt;/p&gt;

&lt;p&gt;LLM marketing materials love to phrase their accuracy in terms of a
&lt;em&gt;percentage&lt;/em&gt;.  Accuracy claims for LLMs in general tend to hover around
70%&lt;sup id=fnref:20:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:20:futzing-fraction-2025-8 id=fnref:20&gt;20&lt;/a&gt;&lt;/sup&gt;.  But these scores vary per field, and when you aggregate them across
multiple topic areas, they start to trend down. This is exactly why “agentic”
approaches for more immediately-verifiable LLM outputs (with checks like “did
the code work”) got popular in the first place: you need to try more than once.&lt;/p&gt;
&lt;p&gt;Independently measured claims about accuracy tend to be quite a bit lower&lt;sup id=fnref:21:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:21:futzing-fraction-2025-8 id=fnref:21&gt;21&lt;/a&gt;&lt;/sup&gt;.
The field of AI benchmarks is exploding, but it probably goes without saying
that LLM vendors game those benchmarks&lt;sup id=fnref:22:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:22:futzing-fraction-2025-8 id=fnref:22&gt;22&lt;/a&gt;&lt;/sup&gt;, because of course every incentive
would encourage them to do that.  Regardless of what their arbitrary scoring on
some benchmark might say, all that matters to &lt;em&gt;your&lt;/em&gt; business is whether it is
accurate for the problems &lt;em&gt;you&lt;/em&gt; are solving, for the way that &lt;em&gt;you&lt;/em&gt; use it.
Which is not necessarily going to correspond to any benchmark. You will need to
measure it for yourself.&lt;/p&gt;
&lt;p&gt;With that goal in mind, our formulation of &lt;math&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;/math&gt; must be a
somewhat harsher standard than “accuracy”.  It’s not merely “was the factual
information contained in any generated output accurate”, but, “is the output
good enough that some given real knowledge-work task is &lt;em&gt;done&lt;/em&gt; and the human
does not need to issue another prompt”?&lt;/p&gt;
&lt;h3 id=surprisingly-small-space-for-slip-ups&gt;Surprisingly Small Space for Slip-Ups&lt;/h3&gt;
&lt;p&gt;The problem with reporting these things as percentages at all, however, is that our actual definition for &lt;math&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;/math&gt; is &lt;math&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mi&gt;attempts&lt;/mi&gt;&lt;/mfrac&gt;&lt;/math&gt;, where &lt;math&gt;&lt;mi&gt;attempts&lt;/mi&gt;&lt;/math&gt; for any given attempt, at least, must be an integer greater than or equal to 1.&lt;/p&gt;
&lt;p&gt;Taken in aggregate, if we succeed on the first prompt more often than not, we &lt;em&gt;could&lt;/em&gt; end up with a &lt;math&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mo&gt;&amp;gt;&lt;/mo&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mfrac&gt;&lt;/math&gt;, but combined with
the previous observation that you almost always have to prompt it more than once, the practical reality is that &lt;math&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;/math&gt; will start at 50% and go down from there.&lt;/p&gt;
&lt;p&gt;If we plug in some numbers, trying to be as &lt;em&gt;extremely&lt;/em&gt; optimistic as we can,
and say that we have a uniform stream of tasks, every one of which can be
addressed by Mallory, every one of which:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;we can measure perfectly, with no overhead&lt;/li&gt;
&lt;li&gt;would take a human 45 minutes&lt;/li&gt;
&lt;li&gt;takes Mallory only a single minute to generate a response&lt;/li&gt;
&lt;li&gt;Mallory will require only 1 re-prompt, so “good enough” half the time&lt;/li&gt;
&lt;li&gt;takes a human only 5 minutes to write a prompt for&lt;/li&gt;
&lt;li&gt;takes a human only 5 minutes to check the result of&lt;/li&gt;
&lt;li&gt;has a per-prompt cost of the equivalent of a single second of a human’s time&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Thought experiments are a dicey basis for reasoning in the face of
disagreements, so I have tried to formulate something here that is absolutely,
comically, over-the-top stacked in favor of the AI optimist here.&lt;/p&gt;
&lt;p&gt;Would that be a profitable?  It sure seems like it, given that we are trading
off 45 minutes of human time for 1 minute of Mallory-time and 10 minutes of
human time.  If we ask Python:&lt;/p&gt;
&lt;div class=highlight&gt;&lt;table class=highlighttable&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class=linenos&gt;&lt;div class=linenodiv&gt;&lt;pre&gt;&lt;span class=normal&gt;1&lt;/span&gt;
&lt;span class=normal&gt;2&lt;/span&gt;
&lt;span class=normal&gt;3&lt;/span&gt;
&lt;span class=normal&gt;4&lt;/span&gt;
&lt;span class=normal&gt;5&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=code&gt;&lt;div&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&amp;gt;&amp;gt;&amp;gt; def FF(H, I, C, P, W, E):
&lt;span class=k&gt;...&lt;/span&gt;     return (W + I + C + E) / (P * H)
&lt;span class=k&gt;...&lt;/span&gt; FF(H=45.0, I=1.0, C=5.0, P=1/2, W=5.0, E=0.01)
...
0.48933333333333334
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;We get a futzing fraction of about 0.4893.  Not bad!  Sounds like, at least
under these conditions, it would indeed be cost-effective to deploy Mallory.
But…  realistically, do you &lt;em&gt;reliably&lt;/em&gt; get useful, done-with-the-task quality
output on the &lt;em&gt;second&lt;/em&gt; prompt?  Let’s bump up the denominator on &lt;math&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;/math&gt; just a little bit there, and see how we fare:&lt;/p&gt;
&lt;div class=highlight&gt;&lt;table class=highlighttable&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class=linenos&gt;&lt;div class=linenodiv&gt;&lt;pre&gt;&lt;span class=normal&gt;1&lt;/span&gt;
&lt;span class=normal&gt;2&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=code&gt;&lt;div&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&amp;gt;&amp;gt;&amp;gt; FF(H=45.0, I=1.0, C=5.0, P=1/3, W=5.0, E=0.01)
0.734
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Oof.  Still cost-effective at 0.734, but not quite as good.  Where do
we cap out, exactly?&lt;/p&gt;
&lt;div class=highlight&gt;&lt;table class=highlighttable&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td class=linenos&gt;&lt;div class=linenodiv&gt;&lt;pre&gt;&lt;span class=normal&gt;1&lt;/span&gt;
&lt;span class=normal&gt;2&lt;/span&gt;
&lt;span class=normal&gt;3&lt;/span&gt;
&lt;span class=normal&gt;4&lt;/span&gt;
&lt;span class=normal&gt;5&lt;/span&gt;
&lt;span class=normal&gt;6&lt;/span&gt;
&lt;span class=normal&gt;7&lt;/span&gt;
&lt;span class=normal&gt;8&lt;/span&gt;
&lt;span class=normal&gt;9&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=code&gt;&lt;div&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class=o&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=kn&gt;from&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=nn&gt;itertools&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=kn&gt;import&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=n&gt;count&lt;/span&gt;
&lt;span class=o&gt;...&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=k&gt;for&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=n&gt;A&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=ow&gt;in&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=n&gt;count&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=n&gt;start&lt;/span&gt;&lt;span class=o&gt;=&lt;/span&gt;&lt;span class=mi&gt;4&lt;/span&gt;&lt;span class=p&gt;):&lt;/span&gt;
&lt;span class=o&gt;...&lt;/span&gt;&lt;span class=w&gt;     &lt;/span&gt;&lt;span class=nb&gt;print&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=n&gt;A&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=n&gt;result&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=o&gt;:=&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=n&gt;FF&lt;/span&gt;&lt;span class=p&gt;(&lt;/span&gt;&lt;span class=n&gt;H&lt;/span&gt;&lt;span class=o&gt;=&lt;/span&gt;&lt;span class=mf&gt;45.0&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=n&gt;I&lt;/span&gt;&lt;span class=o&gt;=&lt;/span&gt;&lt;span class=mf&gt;1.0&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=n&gt;C&lt;/span&gt;&lt;span class=o&gt;=&lt;/span&gt;&lt;span class=mf&gt;5.0&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=n&gt;P&lt;/span&gt;&lt;span class=o&gt;=&lt;/span&gt;&lt;span class=mi&gt;1&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=o&gt;/&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=n&gt;A&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=n&gt;W&lt;/span&gt;&lt;span class=o&gt;=&lt;/span&gt;&lt;span class=mf&gt;5.0&lt;/span&gt;&lt;span class=p&gt;,&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=n&gt;E&lt;/span&gt;&lt;span class=o&gt;=&lt;/span&gt;&lt;span class=mi&gt;1&lt;/span&gt;&lt;span class=o&gt;/&lt;/span&gt;&lt;span class=mf&gt;60.&lt;/span&gt;&lt;span class=p&gt;))&lt;/span&gt;
&lt;span class=o&gt;...&lt;/span&gt;&lt;span class=w&gt;     &lt;/span&gt;&lt;span class=k&gt;if&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=n&gt;result&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=o&gt;&amp;gt;&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=mi&gt;1&lt;/span&gt;&lt;span class=p&gt;:&lt;/span&gt;
&lt;span class=o&gt;...&lt;/span&gt;&lt;span class=w&gt;         &lt;/span&gt;&lt;span class=k&gt;break&lt;/span&gt;
&lt;span class=o&gt;...&lt;/span&gt;
&lt;span class=mi&gt;4&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=mf&gt;0.9792592592592594&lt;/span&gt;
&lt;span class=mi&gt;5&lt;/span&gt;&lt;span class=w&gt; &lt;/span&gt;&lt;span class=mf&gt;1.224074074074074&lt;/span&gt;
&lt;span class=o&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;With this little test, we can see that at our next iteration we are already at
0.9792, and by 5 tries per prompt, even in this absolute fever-dream of an
over-optimistic scenario, with a futzing fraction of 1.2240, Mallory is now a
net detriment to our bottom line.&lt;/p&gt;
&lt;h2 id=harm-to-the-humans&gt;Harm to the Humans&lt;/h2&gt;
&lt;p&gt;We are treating &lt;math&gt;&lt;mi&gt;H&lt;/mi&gt;&lt;/math&gt; as functionally constant so far, an
average around some hypothetical Gaussian distribution, but the distribution
itself can also change over time.&lt;/p&gt;
&lt;p&gt;Formally speaking, an increase to &lt;math&gt;&lt;mi&gt;H&lt;/mi&gt;&lt;/math&gt; would be &lt;em&gt;good&lt;/em&gt; for
our fraction.  Maybe it would even be a good thing; it could mean we’re taking
on harder and harder tasks due to the superpowers that Mallory has given us.&lt;/p&gt;
&lt;p&gt;But an observed increase to &lt;math&gt;&lt;mi&gt;H&lt;/mi&gt;&lt;/math&gt; would probably &lt;em&gt;not&lt;/em&gt; be
good.  An increase could also mean your humans are getting worse at solving
problems, because using Mallory has atrophied their skills&lt;sup id=fnref:23:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:23:futzing-fraction-2025-8 id=fnref:23&gt;23&lt;/a&gt;&lt;/sup&gt; and sabotaged
learning opportunities&lt;sup id=fnref:24:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:24:futzing-fraction-2025-8 id=fnref:24&gt;24&lt;/a&gt;&lt;/sup&gt;&lt;sup id=fnref:25:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:25:futzing-fraction-2025-8 id=fnref:25&gt;25&lt;/a&gt;&lt;/sup&gt;.  It could also go up because your senior,
experienced people now hate their jobs&lt;sup id=fnref:26:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:26:futzing-fraction-2025-8 id=fnref:26&gt;26&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;For some more vulnerable folks, Mallory might just take a shortcut to all these
complex interactions and drive them completely insane&lt;sup id=fnref:27:futzing-fraction-2025-8&gt;&lt;a class=footnote-ref href=#fn:27:futzing-fraction-2025-8 id=fnref:27&gt;27&lt;/a&gt;&lt;/sup&gt; directly.  Employees
experiencing an intense psychotic episode are famously less productive than
those who are not.&lt;/p&gt;
&lt;p&gt;This could all be very bad, if our futzing fraction eventually does head north
of 1 and you need to reconsider introducing human-only workflows, without
Mallory.&lt;/p&gt;
&lt;h1 id=abridging-the-artificial-arithmetic-alliteratively&gt;Abridging the Artificial Arithmetic (Alliteratively)&lt;/h1&gt;
&lt;p&gt;To reiterate, I have proposed this fraction:&lt;/p&gt;
&lt;div style="font-size: 30px; text-align: center;"&gt;
&lt;math&gt;
    &lt;mi&gt;FF&lt;/mi&gt; &lt;mo&gt; = &lt;/mo&gt;
    &lt;mfrac&gt;
        &lt;mrow&gt;&lt;mi&gt;W&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;I&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;
        &lt;mrow&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;!--&lt;mo&gt;✕&lt;/mo&gt;--&gt; &lt;mi&gt;H&lt;/mi&gt;&lt;/mrow&gt;
    &lt;/mfrac&gt;
&lt;/math&gt;
&lt;/div&gt;

&lt;p&gt;which shows us positive ROI when FF is less than 1, and negative ROI when it is
more than 1.&lt;/p&gt;
&lt;p&gt;This model is heavily simplified.  A comprehensive measurement program that
tests the efficacy of &lt;em&gt;any&lt;/em&gt; technology, let alone one as complex and rapidly
changing as LLMs, is more complex than could be captured in a single blog post.&lt;/p&gt;
&lt;p&gt;Real-world work might be insufficiently uniform to fit into a closed-form
solution like this.  Perhaps an iterated simulation with variables based on the
range of values seem from your team’s metrics would give better results.&lt;/p&gt;
&lt;p&gt;However, in this post, I want to illustrate that if you are going to try to
evaluate an LLM-based tool, you need to at &lt;em&gt;least&lt;/em&gt; include some representation
of each of these terms &lt;em&gt;somewhere&lt;/em&gt;.  They are all fundamental to the way the
technology works, and if you’re not measuring them somehow, then you are flying
blind into the genAI storm.&lt;/p&gt;
&lt;p&gt;I also hope to show that a lot of existing assumptions about how benefits
might be demonstrated, for example with user surveys about general impressions,
or by evaluating artificial benchmark scores, are deeply flawed.&lt;/p&gt;
&lt;p&gt;Even making what I consider to be wildly, unrealistically optimistic
assumptions about these measurements, I hope I’ve shown:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;in the numerator, &lt;math&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;/math&gt; might be a lot higher than you
   expect,&lt;/li&gt;
&lt;li&gt;in the denominator, &lt;math&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;/math&gt; might be a lot lower than you
   expect,&lt;/li&gt;
&lt;li&gt;repeated use of an LLM might make &lt;math&gt;&lt;mi&gt;H&lt;/mi&gt;&lt;/math&gt; go up, but despite
   the fact that it's in the denominator, that will ultimately be quite bad for
   your business.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Personally, I don’t have all that many concerns about &lt;math&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/math&gt; and  &lt;math&gt;&lt;mi&gt;I&lt;/mi&gt;&lt;/math&gt;.  &lt;math&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/math&gt; is still seeing significant loss-leader pricing, and &lt;math&gt;&lt;mi&gt;I&lt;/mi&gt;&lt;/math&gt; might not be coming down as fast as vendors would like us to believe, if the other numbers work out I don’t think they make a huge difference.  However, there might still be surprises lurking in there, and if you want to rationally evaluate the effectiveness of a model, you need to be able to measure them and incorporate them as well.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;In particular&lt;/em&gt;, I really want to stress the importance of the influence of LLMs on your &lt;em&gt;team dynamic&lt;/em&gt;, as that can cause massive, hidden increases to &lt;math&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;/math&gt;.  LLMs present opportunities for junior employees to generate an endless stream of chaff that will simultaneously:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;wreck your performance review process by making them look much more
  productive than they are,&lt;/li&gt;
&lt;li&gt;increase stress and load on senior employees who need to clean up unforeseen
  messes created by their LLM output,&lt;/li&gt;
&lt;li&gt;and ruin their own opportunities for career development by skipping over
  learning opportunities.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you’ve already deployed LLM tooling without measuring these things and
without updating your performance management processes to account for the
strange distortions that these tools make possible, your Futzing Fraction may
be much, much greater than 1, creating hidden costs and technical debt that
your organization will not notice until a lot of damage has already been done.&lt;/p&gt;
&lt;p&gt;If you got all the way here, &lt;em&gt;particularly&lt;/em&gt; if you’re someone who is
enthusiastic about these technologies, thank you for reading.  I appreciate
your attention and I am hopeful that if we can start paying attention to these
details, perhaps we can &lt;em&gt;all&lt;/em&gt; stop futzing around so much with this stuff and
get back to doing real work.&lt;/p&gt;
&lt;h2 id=acknowledgments&gt;Acknowledgments&lt;/h2&gt;
&lt;p class=update-note&gt;Thank you to &lt;a href="/pages/patrons.html"&gt;my patrons&lt;/a&gt; who are supporting my writing on
this blog.  If you like what you’ve read here and you’d
like to read more of it, or you’d like to support my &lt;a href="https://github.com/glyph/"&gt;various open-source
endeavors&lt;/a&gt;, you can &lt;a href="/pages/patrons.html"&gt;support my work as a
sponsor&lt;/a&gt;!&lt;/p&gt;
&lt;div class=footnote&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=fn:1:futzing-fraction-2025-8&gt;
&lt;p id=fn:1&gt;I do not share this optimism, but I want to try &lt;em&gt;very&lt;/em&gt; hard in this
particular piece to &lt;em&gt;take it as a given&lt;/em&gt; that genAI is in fact helpful. &lt;a class=footnote-backref href=#fnref:1:futzing-fraction-2025-8 title="Jump back to footnote 1 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:2:futzing-fraction-2025-8&gt;
&lt;p id=fn:2&gt;If we could have a better prompt on demand via some repeatable and
automatable process, surely we would have used a prompt that got the answer
we wanted in the first place. &lt;a class=footnote-backref href=#fnref:2:futzing-fraction-2025-8 title="Jump back to footnote 2 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:3:futzing-fraction-2025-8&gt;
&lt;p id=fn:3&gt;The software idea of a “&lt;a href="https://www.w3.org/WAI/UA/work/wiki/Definition_of_User_Agent"&gt;user
agent&lt;/a&gt;”
straightforwardly comes from the legal principle of an
&lt;a href="https://en.wikipedia.org/wiki/Law_of_agency"&gt;agent&lt;/a&gt;, which has deep roots
in common law, jurisprudence, &lt;a href="https://en.wikipedia.org/wiki/Principal–agent_problem"&gt;philosophy, and
math&lt;/a&gt;.  When we
think of an agent (some software) acting on behalf of a principal (a human
user), this historical baggage imputes some &lt;a href="https://blog.glyph.im/2005/11/ethics-for-programmers-primum-non.html"&gt;&lt;strong&gt;important ethical
obligations&lt;/strong&gt;&lt;/a&gt;
to the developer of the agent software.  genAI vendors have been as eager
as any software vendor to &lt;a href="https://openai.com/policies/row-terms-of-use/#:~:text=NEITHER%20WE%20NOR%20ANY%20OF%20OUR%20AFFILIATES%20OR%20LICENSORS%20WILL%20BE%20LIABLE"&gt;dodge responsibility for faithfully representing
the user’s
interests&lt;/a&gt;
even as there are some indications that &lt;a href="https://www.forbes.com/sites/marisagarcia/2024/02/19/what-air-canada-lost-in-remarkable-lying-ai-chatbot-case/"&gt;at least some courts are not
persuaded&lt;/a&gt;
by this dodge, at least by the consumers of genAI attempting to pass on the
responsibility all the way to end users.  Perhaps it goes without saying,
but I’ll say it anyway: I don’t like this newer interpretation of “agent”. &lt;a class=footnote-backref href=#fnref:3:futzing-fraction-2025-8 title="Jump back to footnote 3 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:4:futzing-fraction-2025-8&gt;
&lt;p id=fn:4&gt;&lt;a href="https://arxiv.org/abs/2502.15840"&gt;“Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous
Agents”&lt;/a&gt;, Axel Backlund, Lukas Petersson,
Feb 20, 2025 &lt;a class=footnote-backref href=#fnref:4:futzing-fraction-2025-8 title="Jump back to footnote 4 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:5:futzing-fraction-2025-8&gt;
&lt;p id=fn:5&gt;&lt;a href="https://xcancel.com/leojr94_/status/1901560276488511759"&gt;“random thing are happening, maxed out usage on api keys”&lt;/a&gt;, @leojr94 on Twitter, Mar 17, 2025 &lt;a class=footnote-backref href=#fnref:5:futzing-fraction-2025-8 title="Jump back to footnote 5 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:6:futzing-fraction-2025-8&gt;
&lt;p id=fn:6&gt;&lt;a href="https://apnews.com/article/chatgpt-study-harmful-advice-teens-c569cddf28f1f33b36c692428c2191d4"&gt;“New study sheds light on ChatGPT’s alarming interactions with
teens”&lt;/a&gt; &lt;a class=footnote-backref href=#fnref:6:futzing-fraction-2025-8 title="Jump back to footnote 6 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:7:futzing-fraction-2025-8&gt;
&lt;p id=fn:7&gt;&lt;a href="https://apnews.com/article/artificial-intelligence-chatgpt-fake-case-lawyers-d6ae9fa79d0542db9e1455397aef381c"&gt;“Lawyers submitted bogus case law created by ChatGPT. A judge fined
them
$5,000”&lt;/a&gt;,
by Larry Neumeister for the Associated Press, June 22, 2023 &lt;a class=footnote-backref href=#fnref:7:futzing-fraction-2025-8 title="Jump back to footnote 7 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:8:futzing-fraction-2025-8&gt;
&lt;p id=fn:8&gt;During which a human will be busy-waiting on an answer. &lt;a class=footnote-backref href=#fnref:8:futzing-fraction-2025-8 title="Jump back to footnote 8 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:9:futzing-fraction-2025-8&gt;
&lt;p id=fn:9&gt;Given the fluctuating pricing of these products, and fixed subscription overhead, this will obviously need to be amortized; including all the additional terms to actually convert this from your inputs is left as an exercise for the reader. &lt;a class=footnote-backref href=#fnref:9:futzing-fraction-2025-8 title="Jump back to footnote 9 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:10:futzing-fraction-2025-8&gt;
&lt;p id=fn:10&gt;I feel like I should emphasize explicitly here that everything is an
average over repeated interactions.  For example, you might observe that a
particular LLM has a low probability of outputting acceptable work on the
first prompt, but higher probability on subsequent prompts in the same
context, such that it usually takes 4 prompts.  For the purposes of this
extremely simple closed-form model, we’d still consider that a
&lt;math&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;/math&gt; of 25%, even though a more sophisticated model, or
a monte carlo simulation that sets progressive bounds on the probability,
might produce more accurate values. &lt;a class=footnote-backref href=#fnref:10:futzing-fraction-2025-8 title="Jump back to footnote 10 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:11:futzing-fraction-2025-8&gt;
&lt;p id=fn:11&gt;&lt;a href="https://social.coop/@chrisjrn/115011133688436556"&gt;No it isn’t,
actually&lt;/a&gt;, but for the
sake of argument let’s grant that it is. &lt;a class=footnote-backref href=#fnref:11:futzing-fraction-2025-8 title="Jump back to footnote 11 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:12:futzing-fraction-2025-8&gt;
&lt;p id=fn:12&gt;It’s worth noting that all this expensive measuring &lt;em&gt;itself&lt;/em&gt; must be
included in &lt;math&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;/math&gt; until you have a solid grounding for
all your metrics, but let’s optimistically leave all of that out for the
sake of simplicity. &lt;a class=footnote-backref href=#fnref:12:futzing-fraction-2025-8 title="Jump back to footnote 12 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:13:futzing-fraction-2025-8&gt;
&lt;p id=fn:13&gt;&lt;a href="https://www.newsweek.com/nearly-half-employees-trust-ai-more-their-coworkers-2113159"&gt;“AI Company Poll Finds 45% of Workers Trust the Tech More Than Their
Peers”&lt;/a&gt;,
by Suzanne Blake for Newsweek, Aug 13, 2025 &lt;a class=footnote-backref href=#fnref:13:futzing-fraction-2025-8 title="Jump back to footnote 13 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:14:futzing-fraction-2025-8&gt;
&lt;p id=fn:14&gt;&lt;a href="https://www.cmu.edu/dietrich/news/news-stories/2025/july/trent-cash-ai-overconfidence.html"&gt;AI Chatbots Remain Overconfident — Even When They’re
Wrong&lt;/a&gt;
by Jason Bittel for the Dietrich College of Humanities and Social Sciences
at Carnegie Mellon University, July 22, 2025 &lt;a class=footnote-backref href=#fnref:14:futzing-fraction-2025-8 title="Jump back to footnote 14 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:15:futzing-fraction-2025-8&gt;
&lt;p id=fn:15&gt;&lt;a href="https://spectrum.ieee.org/ai-mistakes-schneier"&gt;AI Mistakes Are Very Different From Human
Mistakes&lt;/a&gt; by Bruce Schneier
and Nathan E. Sanders for IEEE Spectrum, Jan 13, 2025 &lt;a class=footnote-backref href=#fnref:15:futzing-fraction-2025-8 title="Jump back to footnote 15 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:16:futzing-fraction-2025-8&gt;
&lt;p id=fn:16&gt;Foreshadowing is a narrative device in which a storyteller gives an
advance hint of an upcoming event later in the story. &lt;a class=footnote-backref href=#fnref:16:futzing-fraction-2025-8 title="Jump back to footnote 16 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:17:futzing-fraction-2025-8&gt;
&lt;p id=fn:17&gt;&lt;a href="https://www.ipsos.com/en-us/people-are-worried-about-misuse-ai-they-trust-it-more-humans"&gt;“People are worried about the misuse of AI, but they trust it more than humans”&lt;/a&gt; &lt;a class=footnote-backref href=#fnref:17:futzing-fraction-2025-8 title="Jump back to footnote 17 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:18:futzing-fraction-2025-8&gt;
&lt;p id=fn:18&gt;&lt;a href="https://youtu.be/w3EZpcTZ4ZA?si=816uBg6N3Pmon2P3"&gt;“Why I stopped using AI (as a Senior Software
Engineer)”&lt;/a&gt;, theSeniorDev
YouTube channel, Jun 17, 2025 &lt;a class=footnote-backref href=#fnref:18:futzing-fraction-2025-8 title="Jump back to footnote 18 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:19:futzing-fraction-2025-8&gt;
&lt;p id=fn:19&gt;&lt;a href="https://www.youtube.com/watch?v=1ghBG302_LQ"&gt;“I was an AI evangelist. Now I’m an AI vegan. Here’s
why.”&lt;/a&gt;, Joe McKay for the
greatchatlinkedin YouTube channel, Aug 8, 2025 &lt;a class=footnote-backref href=#fnref:19:futzing-fraction-2025-8 title="Jump back to footnote 19 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:20:futzing-fraction-2025-8&gt;
&lt;p id=fn:20&gt;&lt;a href="https://originality.ai/blog/what-llm-is-the-most-accurate"&gt;“What LLM is The Most Accurate?”&lt;/a&gt; &lt;a class=footnote-backref href=#fnref:20:futzing-fraction-2025-8 title="Jump back to footnote 20 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:21:futzing-fraction-2025-8&gt;
&lt;p id=fn:21&gt;&lt;a href="https://futurism.com/the-byte/study-chatgpt-answers-wrong"&gt;“Study Finds That 52 Percent Of ChatGPT Answers to Programming Questions are Wrong”&lt;/a&gt;, by Sharon Adarlo for Futurism, May 23, 2024 &lt;a class=footnote-backref href=#fnref:21:futzing-fraction-2025-8 title="Jump back to footnote 21 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:22:futzing-fraction-2025-8&gt;
&lt;p id=fn:22&gt;&lt;a href="https://blog.boxcars.ai/p/off-the-mark-the-pitfalls-of-metrics"&gt;“Off the Mark: The Pitfalls of Metrics Gaming in AI Progress
Races”&lt;/a&gt;, by
Tabrez Syed on BoxCars AI, Dec 14, 2023 &lt;a class=footnote-backref href=#fnref:22:futzing-fraction-2025-8 title="Jump back to footnote 22 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:23:futzing-fraction-2025-8&gt;
&lt;p id=fn:23&gt;&lt;a href="https://thomasorus.com/i-tried-coding-with-ai-i-became-lazy-and-stupid"&gt;“I tried coding with AI, I became lazy and
stupid”&lt;/a&gt;,
by Thomasorus, Aug 8, 2025 &lt;a class=footnote-backref href=#fnref:23:futzing-fraction-2025-8 title="Jump back to footnote 23 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:24:futzing-fraction-2025-8&gt;
&lt;p id=fn:24&gt;&lt;a href="https://www.psychologytoday.com/us/blog/the-algorithmic-mind/202505/how-ai-changes-student-thinking-the-hidden-cognitive-risks"&gt;“How AI Changes Student Thinking: The Hidden Cognitive
Risks”&lt;/a&gt;
by Timothy Cook for Psychology Today, May 10, 2025 &lt;a class=footnote-backref href=#fnref:24:futzing-fraction-2025-8 title="Jump back to footnote 24 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:25:futzing-fraction-2025-8&gt;
&lt;p id=fn:25&gt;&lt;a href="https://phys.org/news/2025-01-ai-linked-eroding-critical-skills.html"&gt;“Increased AI use linked to eroding critical thinking skills”&lt;/a&gt; by Justin Jackson for Phys.org, Jan 13, 2025 &lt;a class=footnote-backref href=#fnref:25:futzing-fraction-2025-8 title="Jump back to footnote 25 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:26:futzing-fraction-2025-8&gt;
&lt;p id=fn:26&gt;&lt;a href="https://dev.to/manuartero/ai-could-end-my-job-just-not-the-way-i-expected-5g3m"&gt;“AI could end my job — Just not the way I expected”&lt;/a&gt; by Manuel Artero Anguita on dev.to, Jan 27, 2025 &lt;a class=footnote-backref href=#fnref:26:futzing-fraction-2025-8 title="Jump back to footnote 26 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=fn:27:futzing-fraction-2025-8&gt;
&lt;p id=fn:27&gt;&lt;a href="https://www.psychologytoday.com/us/blog/urban-survival/202507/the-emerging-problem-of-ai-psychosis"&gt;“The Emerging Problem of “AI
Psychosis””&lt;/a&gt;
by Gary Drevitch for Psychology Today, July 21, 2025. &lt;a class=footnote-backref href=#fnref:27:futzing-fraction-2025-8 title="Jump back to footnote 27 in the text"&gt;↩&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;&lt;/body&gt;</content><category term="misc"></category><category term="ai"></category><category term="llm"></category><category term="basic-arithmetic"></category></entry></feed>