Load-Bearing - Vocabulary trend analysis of Claude coding agents

Show HN: The load-bearing vocabulary of Claude

Load-Bearing is a data-driven tool that analyzes vocabulary trends in GitHub pull requests to reveal the distinctive language of AI coding agents like Claude. By clustering over 47,000 PRs into 8 vocabulary clusters, it identifies a dominant cluster representing 45% of human-attributed PRs, characterized by words like 'load-bearing' and 'seamlessly'. This tool offers unique insights into how AI influences coding communication, making it valuable for developers, researchers, and AI enthusiasts.

We scraped 100 GitHub PRs daily to uncover the load-bearing vocabulary of Claude, revealing a new linguistic fingerprint in coding.
  1. nater5000

    I was pleasantly surprised when I attempted to scroll down and realized everything the author wanted to present fit on-screen. It's almost ironic that this site is able to make such an obvious, compelling presentation without being overly verbose or complicated (something which LLMs have a hard time doing). I wouldn't read TOO deeply into what is being presented, but the author has done a good job to not inject their own bias into the presentation which works well.

    I suspect, as we continue forward, humans will slowly start to adopt the language of LLMs, or at least certain language quirks that come from interacting with LLMs. Something I've noticed in my own writing is that I now present lists of examples in a consistent way: "... such as <example 1>, <example 2>, etc., ...". I started to notice I was using this pattern quite a bit somewhat recently, but I took a quick look at some of my social media posts and realized it's been occurring for a while. I had realized that I grown accustomed to this kind of language because, especially early on, LLMs would focus too much on the specific examples I'd provide when, really, I was just trying to give them a sense of what I was looking for. I just picked up that providing two examples then adding the "etc." worked to get the LLM to not focus so much on the specific examples and to understand that they need to consider more than what I explicitly presented. Of course, now I write like that in my social media comments, in Slack with […]

  2. ben30

    Interesting response from Claude, when I attempted to reduce the "load-bearing" in its responses:

    I added to my global prompt:

    - Orwell's first rule: never use a metaphor you're used to seeing in print. "Load-bearing", "the crux", "first-class citizen" signal insight instead of showing it. Name the specific mechanism

    when I asked what it thought of the change, its reply was:

    The Orwell bullet fights my own system prompt. My harness instructions literally tell me to flag "something load-bearing" when I find it.

  3. Labo333

    Author here! Grateful for the kind words, human communities like HN really hit differently when you spend the whole day chatting with sycophantic and bullshitting agents (including to make this page).

    I'm currently adding a search bar as well as increasing the data to 1000 PR per day.

    A nice thing that is not obvious on the main page is that the dataset and analysis are updated daily using Github Actions (at least when they don't suffer from an outage ^^). I find it pretty cool to be able to build such apps without a "backend"!

  4. SalariedSlave

    I've recently seen this mentioned more and more, both on HN and on reddit. It seems these output patterns are getting worse. It's not just Claude, my impression is that all of the current models have this style issue. Their writing can get borderline incomprehensible.

    Is there some feedback loop or compounding happening with each model generation?

    Maybe newer models are ingesting too much AI content?

    If the ratio of AI generated content in training data is getting higher and higher (because the amount of AI generated content is increasing in general), maybe this is a compounding bias, poisoning the training?

  5. prmph

    Yeah, Claude's writing is weird nowadays, and it uses words that are weird and non-standard in context, and its sentence structure is sometimes weird as well.

    Working with it on my code, it's now frequently making weird word choices such as:

    - "name(s)" as a verb (instead of "specify/specifies", etc), e.g., "...the function names the argument"

    - "carries" instead of "contains"

    - "verdict" instead of "result"

    - "judge", instead of "validate"

    Very weird. Also, in writing comments and docs, it is terse in ways that make the writing difficult to parse, like omitting mentioning what a noun refers to, e.g. abc does not accept..." instead of "The abc function does not accept..."

    I've resorted to banning from using certain words, and keep asking it to rewrite its text more clearly.

  6. sosull

    I really love this. It’s comprehensive, it consolidates the data to the point where the argument effectively ‘makes itself’, and the way it’s presented respects the reader’s time. It also makes for an interesting challenge (for me at least) to try to characterise the subject matter of a language problem so narrowly.

    No ream of slides. No narrative. Just a lovely big painful conclusion.

  7. Jordan-117

    I wonder to what extent this is the result of suboptimal RLHF versus the inherent intelligence of the model making its language more intricate and difficult for humans to easily parse? On the one hand, it's a common trope that highly educated people can talk in a way that's confusing and annoying to regular people who don't know all the jargon. But on the other hand, it's a mark of a skilled communicator to be able to efficiently distill complex information to its bare essentials in an easily-digestible way. Of course, that also seems to imply that these models are working at a higher level and need to talk down to us to an extent. Or maybe "Claudish" is just akin to stuff like "caveman", raw chain of thought, neuralese, etc., which are likewise much more dense/efficient but harder to interpret?

  8. stabbles

    Besides vocabulary I'd also be curious about how it structures sentences.

    The "X, not Y" is well known, but another thing that bothers me is "It <verb>s no <noun>" instead of "It doesn't <verb> <noun>".

    For example: "the list contains no string" or "it changes no behavior" or "it holds no directory".

  9. loglog

    As of now, there are only 2 natural language clusters in the partition: Claudish and Spanish/French. From a first glance, the other clusters seem centered on technologies (lots of command line flags and abbreviations among the "most representative" words). If the goal is to investigate English language usage trends, it might be better to cluster based only on English words.

  10. wavewrangler

    Was talking about the use of shipped recently, and I was mocked for asking such a crazy question, by freshly self-minted engineers, no less. No wonder they thought it was ridiculous...it had been a part of their vocabulary their entire career. All few weeks of it. I wonder what those guys are doing now. This was about a month ago. Do you think what they shipped ever...landed?

More from this day

2026-08-27