Tomas Vondra

Tomas Vondra

blog about Postgres code and community

Are we reverting patches because of bugs found by AI?

Over the past month, several major features got reverted from Postgres 19. At this point, even if everything else goes perfectly, Postgres 19 will be released about a month later than originally planned. A couple recent blog posts seem to suggest the reverts happened due to AI finding complex bugs in those patches, with fixes way too invasive this late in the development cycle. I don’t think that narrative matches reality, and I think it matters to correct that.

For example, a recent blog post by Elizabeth Christensen says:

The AI factor

AI has to be part of the story for why we’re in this situation. Several folks are using AI tools to find bugs and create valid reproducible test cases. The fixes are fairly large, so pushing to a future version makes the most sense for many of these features.

If you want an example of how AI is changing PostgreSQL’s deployment pipeline, think about CVE patching. PostgreSQL used to average a couple CVEs each release; the August 2026 patch for Postgres 18 had 28 CVEs.

PostgreSQL currently has no official AI contribution policy, though one is expected soon.

This clearly implies the reverts are happening due to AI finding bugs that are too complex for this release, and so the features get reverted. Various other blog posts make similar claims, linking to this post as a source.

I don’t think that narrative is quite accurate. I’ve been involved in some of the patches in some way, and that claim simply does not match my understanding of why the reverts happened.

And I think it matters. We are clearly facing some challenges, otherwise the release would not slip by a month (after many years when we managed to get releases out on time). And if we misidentify the reasons, it’s hard to address the actual causes.

Note: I have nothing against Elizabeth, I like her blog posts and I’ve even been a guest on the PgUS meetup she’s running (and it was fun). It’s not personal, I’m merely responding to a blog post.

Patches reverted from 19

I went through the threads that led to the recent reverts, and tried to identify if the issues came from an AI or not. Here’s a brief summary of what I found. I’ll present the main reverts in reverse order by date, with a brief summary of what led to the revert / if AI was the reason.

Each revert commit has a link to the relevant mailing list thread, with all the details. I encourage others to verify my conclusions. A lot of this is my personal opinion, not hard data.

There’s also a fair amount of discussion in the scary patch contest thread, which likely forced some of the difficult decisions.

Of course, it’s hard to know for sure - maybe some AI tool was used but not disclosed. (But then probably still works against the AI claim, no?)

2026-09-16 / Revert online data checksum transitions

  • revert: c05d5ce1236

  • None of the issues reported earlier seems to have been found by AI.

  • There is a thread about breaking online checksums with Claude. But that’s multiple days after Robert’s “scary patch” message mentioning this patch. I would not rate any of the AI findings as “critical” of “unfixable before release”.

  • Overall, I think the revert was simply a consequence of there being too many post-commit issues, some of which were found fairly late, and being cautious and mitigating risk.

  • conclusion: not AI

2026-09-15 / Revert UPDATE/DELETE FOR PORTION OF

  • revert: a9d2f728240

  • The review that ultimately led to the revert came from Andres, and it’s clearly a human-written review.

  • There’s a number of issues, most importantly an incorrect behavior under concurrency - which was even known to the authors, and explained in the comments / docs with a “workaround” (but as Andres pointed out, that’s not really acceptable).

  • conclusion: not AI

2026-09-13 / Revert pg_get_role_ddl(), pg_get_tablespace_ddl(), and pg_get_database_ddl().

  • revert: db169985c10

  • The revert discussion started with Noah’s human review, pointing out issues with dependencies and a couple other things (pg_dump differences etc.)

  • He even mentions he did review using Opus 4.8, but that it found just minor / less significant issues.

  • The thread then spun into a much deeper discussion about the design, code duplication, etc.

  • Andres also posted his human review with yet more questions/concerns about permissions, locking, etc.

  • conclusion: not AI

2026-09-11 / Revert “Support more object types within CREATE SCHEMA”

  • revert(s): e0fdc3f54b4, 3c5d28ba64e

  • This seems to actually originate in an Opus 5 review, which identified some issues with the patch changing semantics established in PG18.

  • conclusion: AI

2026-09-10 / Remove batching from RI fast-path checks

  • revert: 25649d6e791

  • The feature was considered “concerning” due to many post-commit fixes, suggesting it may not be mature enough.

  • I found one LLM review by Nikolay, but that’s 3 months before the revert, and seems to have been addressed.

  • Ultimately, what got it reverted were issues with performing FK checks in the right order with batching. Some of the issues/reproducers seem to have been found by AI, but not sure.

  • conclusion: possibly AI

2026-09-07 / Revert SQL Property Graph Queries (SQL/PGQ)

  • revert: 2b9e1aff4d3

  • The patch was considered scary, both due to the size of the patch, how many parts of the code base it touched (parser, rewriter, executor, …) and the number of post-commit fixes (about 1/3 of the reverted commits are “Fix something”).

  • What ultimately led to the review was this human review by Andres, followed by another one.

  • Andres also did an AI review, but AFAICS that mostly just added up to the earlier findings.

  • conclusion: not AI

2026-08-27 / Revert support for ALTER TABLE … MERGE/SPLIT PARTITION(S) commands

  • revert: 3e8bcc8644f

  • The revert thread was started by a human review by Zsolt. I believe it was a regular review.

  • There’s some Claude use later in the thread, but it seems mostly immaterial for the outcome.

  • This is clearly a hard feature, it was already committed / reverted in the PG17 cycle (3890d90c150).

  • conclusion: not AI

2026-07-17 / Revert “Add GROUP BY ALL”.

  • revert: 372b8d1adb7

  • Has bugs that would require refactoring too invasive this late in the cycle.

  • The report does not mention any AI tools.

  • conclusion: probably not AI

2026-06-18 / Revert non-text output formats for pg_dumpall

  • revert: 7ca548f23a6

  • Seemed not ready, based on this human review by Noah.

  • Noah even mentions many of the points were identified in PG18 cycle, so very unlikely to be from AI review.

  • conclusion: not AI

2026-06-08 / Revert “Enable fast default for domains with non-volatile constraints”

  • revert: a0354e29c41

  • Incomplete design, needs TAM changes (of scope post feature freeze)

  • The report does not mention any AI tools.

  • conclusion: probably not AI

2026-05-23 / Revert “Allow logical replication snapshots to be database-specific”

2026-05-20 / Revert “Reject degenerate SPLIT PARTITION with DEFAULT partition”

  • revert: 0392fb900eb

  • per buildfarm failures

  • conclusion: not AI

To sum this up, from the 12 reverts, only 4 came out at “probably AI” or “AI”. The remaining reverts seem to have happened due to issues identified by regular human review.

I don’t see how this could be interpreted as “AI is finding bugs so complex we have to revert features.” That interpretation could even be a little bit insulting to the reviewers who actually found the issues.

Reverts in the past

I’ve been wondering how we’re doing compared to previous releases - are we reverting more features, or are there other differences. Some blog posts suggest the number of reverts is unprecedented or “crazy”. Is it?

I looked at git history, and counted the number of reverts for releases starting at Postgres 14. Here’s a chart with the number of reverts by day of the development cycle:

We usually fork the next development branch at the end of June of the previous year, so for example Postgres 19 was forked on June 30, 2025. This counts as “day 0” in the chart. Then development happens, until the “feature freeze” at the beginning of April. That’s day ~275. And then the code stabilization happens - testing, fixing bugs etc. Releases usually happen on the day ~480.

I’ve included a bit more time to show post-release reverts (but we don’t have very many of those, and it tends to be small things).

Note: This simply counts commits with subjects starting with “Revert”. That might miss a couple of reverts where the committer changed the subject to something like “Remove XYZ”, but that’s rare. Of course, it’s still just a number of commits, it does not say how large the reverts are etc.

By this chart it doesn’t seem like we’re doing much worse than in previous years. We’re about on track with past releases, except for PG 16 which seems to have had many fewer reverts.

The development cycles seem to be comparable, with about the same number of commits (between forking and feature freeze). PG14 has about 1700 commits, PG19 has about 2200. Not a huge difference, and it’s favorable for PG19 (the fraction of reverted commits gets lower).

Of course, “number of commits” is a very crude metric, and it may not be a great measure of development activity. A bit like “lines of code” are not a great measure of productivity.

What clearly changed is the timing of the reverts. In previous years, most of the reverts happened right after feature freeze, around day 300. For PG17 it’s particularly clear, but other releases behave similarly.

For PG19 the activity simply flatlined. There’s not even a small bump! Then on day ~430 it picks up - that’s the series of reverts at the beginning of September.

Here’s a chart with “size of reverts”, measured by “insertions” and “deletions” in the diffs (using git show --stat).

The story is about the same - we’ve been doing about the same as in previous years until feature freeze, then nothing happens until the spike at the beginning of September. Some of the reverted features were clearly pretty large.

While I don’t think the reverts are due to bugs found by AI, that does not mean AI is not part of the story. As Elizabeth pointed out in the part I quoted earlier:

If you want an example of how AI is changing PostgreSQL’s deployment pipeline, think about CVE patching. PostgreSQL used to average a couple CVEs each release; the August 2026 patch for Postgres 18 had 28 CVEs.

I believe this is spot on, and I believe it’s also one of the primary reasons for the lower review activity during the stabilization period. A significant number of senior/experienced contributors are busy with handling security reports, some of which come from AI. Last year we had maybe 5 CVEs, this year we’re already at 44. Of course, we have to fix the reported issues, and that takes a fair amount of time.

A couple months ago I wrote about AI inversion and that I’m concerned about the impact it might have on the project. I think this may be a good example of such impact.

It’d be naive to think AI does not have similar effects in other parts of the community, e.g. for regular patches. It’s just harder to quantify, while for security reports we have the list of CVEs.

Do you have feedback on this post? Please reach out by e-mail to tomas@vondra.me.