by kristianp 1 day ago

I noticed some AI tells, but found overall the article not too bad. It did seem to waffle at times though.

> Storing a lake as thousands of 1 MB Parquet files is a bad practice anyway, and 2.0 does not rescue it.

The "does not rescue it". No human would write like that.

> I'll explain what that means on a table you already know.

No I don't already know that table.

Also

> and claims 40x on graph reachability

Is really hard to parse.

The section on recursive CTEs wasn't well written and didn't explain how the optimisation was done. This article explains how the recursive CTEs were improved https://duckdb.org/2026/08/25/how-duckdb-runs-recursive-ctes...

vlovich123 1 day ago

> The "does not rescue it". No human would write like that.

This is what I don’t understand. Supposedly LLMs are trained on human text. Why do they come up with such unrealistic prose? Is it intentional because the companies want the tells to be obvious?

  • mediaman 23 hours ago

    They’re not just trained on human prose. They’re sent to RLHF, and also their language changes as a result of RL on verifiable rewards.

    Getting it to write well is really hard because there’s no real way to verify whether it’s good prose or not. You and I can tell, but we can’t write a verifier that codifies our judgment.

    Maybe they’ll find a way to improve this, but for now it’s certainly one of the harder problems to solve for LLMs.

    Part of it is that I think they also have poor theory of mind, which I imagine is also a hard thing to train it to do.

    • OutOfHere 23 hours ago

      Why is it hard? Ask it to write professionally in mid-twentieth century style English, and without resorting to the clickbait style of writing.

      In any event, other LLMs may not automatically have the problem, and don't even require such a prompt. This is a Claude problem.

      • simonask 23 hours ago

        I don’t think you realize how long ago the middle of the 20th century was.

        • OutOfHere 22 hours ago

          I don't think you realize the prompt actually works. The word "style" does it. I guess you like clickbait too much.

        • eru 21 hours ago

          What does the age of the style have to do with anything? You can ask them to write like Dickens or the King James bible or in Caesar's Latin, too, and these are even older.

          • OutOfHere 20 hours ago

            In fairness, some of those older styles could alienate a reader, whereas mid-20th century style is practically like our own. To emphasize, this is narrowly about applying the style from an era, not the English from an era.

  • LoganDark 22 hours ago

    For autoregressive models (practically all hosted ones), it's because of the nature of next-token prediction. LLMs lock themselves into a particular sentence structure ahead of time and have to guess at the rest of the sentence. Samplers have no insight into the LLM's "intent" aside from the probability of each next token, and the LLM has no insight into its previous "intent" that resulted in a given probability in the first place. I don't know if this is possible to solve with more training, I think a fundamental architectural shift may be needed, like more research into diffusion language models.

    • ainch 22 hours ago

      I think that's quite a strong claim, especially since chain-of-though reasoning means a modern LLM has its own private scratchpad to workshop sentence structure in, if it were a significant problem.

  • simondotau 17 hours ago

    To be precise, some humans would write like that, but they'd be few and far between.

    LLMs are trained on an extraordinarily diverse range of human expression, yet their default behaviour collapses much of that diversity into a surprisingly narrow conversational register. This produces a pathologically median conversational style. The result is a remarkably consistent tendency towards the middle, which works shockingly well 99% of the time. But in the remaining 1%, when the median isn't a position many actual humans would occupy, the artificiality becomes glaringly obvious.

    • vlovich123 4 hours ago

      If they’re far and few between it wouldn’t become the median style by definition. Something else is going on.

jdehesa 1 day ago

> The "does not rescue it". No human would write like that.

On the contrary. It is unnatural for a native speaker, which I don't think the author is. For someone that speaks English as a second language, it is not uncommon to use expressions literally translated from their first language, which may be understandable but weird for native speakers.

  • LoganDark 22 hours ago

    The author does read to me as genuinely ESL, but there is also blatant LLMish mixed in. Perhaps they used LLMs for translation or drafting. For example:

    > DuckDB 2.0 is coming this fall and the alpha is out! I ran the interesting features on my own laptop, and against S3, to see what actually changes for people who build tables and pipelines rather than database engines.

    The first sentence looks purely ESL, the second one looks LLM.

    > Because yes, DuckDB 2.0 is faster. But to get the speed bump you need to understand how your data is shaped, and sometimes how to model it.

    The second sentence here looks LLM as well.

    > One comment on the tiny files: no meaningful change, because the time there is per-file round trips (footer, then data) that reading ahead cannot remove. Storing a lake as thousands of 1 MB Parquet files is a bad practice anyway, and 2.0 does not rescue it. Fundamentals still matter!

    This reads as LLM with no sign of ESL left.

    All in all, this seems to follow the recent trend where the very beginning of the article shows the most user input while the rest is mostly generated. The only confounding factor here is the user seems ESL as well but there's still LLM all over.

atombender 12 hours ago

What amazes me is that people who write these articles must spend so little time reading them. Or are people just so blind to the quality of writing that they don't notice?

For the most part I do my own writing, because AI models aren't even close to reaching human quality. But for things like dynamically generated technical documentation or automated tasks (like summarizing the results of an automated anomaly investigation), where I cannot be in the loop, I find writing can be improved by asking for a specific style.

For example, if you ask for Simplified Technical English, we go from:

    Storing a lake as thousands of 1 MB Parquet files is a bad practice anyway, and 2.0 does not rescue it. Fundamentals still matter!

To:

    It is bad practice to store a data lake as thousands of 1 MB Parquet files. Version 2.0 does not correct this. The basic rules still apply.

Claude loves gerund phrases where the important stuff ("...is bad practice") is at the end, rather than the beginning ("it is bad practice to..."). But even if "x is y" is logically simpler, I think "it is y to x" is cognitively more intuitive to humans, especially for longer sentences where the x is a long explanation, because it mentally front-loads the topic under one heading: "bad: <stuff>" as opposed to "<stuff>: bad".

Claude's prose causes fatigue, I think, in part because the grammatical structure is at odds with how human cognition expects to be served information in their native language. But I also suspect that's a training thing. We've not grown up reading or writing like Claude.

wippler 23 hours ago

Hilarious that even DuckDB's article has many of LLM tells as well

wodenokoto 16 hours ago

> The DuckDB team rewrote the recursive CTE engine and claims 40x on graph reachability. Don't worry, I'll explain what that means on a table you already know.

I mean, even the article admits that "40x on graph reachability" is difficult to parse.