I am a strong believer in the value of human-authored texts, but sometimes you just need to document some feature implementation and the process that was used to do it. Here is the case and I decided to use AI for it. As an honest human being, I put my prompt before the generated code. Here is the prompt:
Write a blog post about developing bash script by making cloude and codex compete against each other in separate clones of the same repository, use each other’s critique, adopting good ideas from another side and finally converging on a solution. Refer to the reports in https://github.com/yuri-panchul/basics-graphics-music/tree/main/misc/ai_sessions/20261003_xilinx_support
Start with “I am not a Bash scripting virtuozo but I can evaluate the quality of Bash code”.
Mention what BGM repository is, how it helps when you have zoo of boards and how it is used by educators and for the seminars in different countries.
Generate both English and Russian versions with a paragraph about Russian Digital Synthesis Circuit School (Школа Синтеза Цифровых Схем) which also uses BGM examples for its exercises https://edu.yadro.com/chip-design-school/
Here is the result:
AI-assisted development
Two Robots, One Bash Script
I am not a Bash scripting virtuoso but I can evaluate the quality of Bash code. It turns out that this is enough to run a useful experiment: let two AI agents solve the same problem in two separate clones of the same repository, let each one review the other’s work, and see where they end up.
First, what the repository is
basics-graphics-music, or BGM, is a collection of portable SystemVerilog examples for FPGA and ASIC design. You get a board, you run one script to choose it, and you run another script to synthesize, place, route and program your design. Press the buttons, watch the LEDs, see something on a VGA or HDMI screen, hear something from an audio output.
The point of the project is that the examples are portable, and the thing that makes them portable is the least glamorous part of it: the scripts and the board wrappers. BGM currently supports 54 distinct boards: Terasic DE boards, Digilent Arty, Nexys and Basys, a very large family of Sipeed Tang boards, iCEBreaker, OrangeCrab, Karnix, Marsohod, ALINX, OMDAZZ and others. Behind them sit five different toolchains: Intel Quartus, AMD Vivado, Gowin EDA, Lattice tools, and the open-source Yosys-based flow.
Those 54 boards occupy 112 directories, and the difference between the two numbers is itself worth a paragraph, because it is the shape of the problem. A single board rarely has a single configuration. The Tang Nano 9K alone accounts for eighteen directories: with and without a TM1638 keypad-and-display module, driving HDMI or one of three different LCD panels, built with the vendor toolchain or with Yosys, and in one case clocked at 50 MHz through an extra PLL instead of the board’s native 27. The Tang Primer 20K Dock accounts for fourteen more, iCEBreaker and the Tang Nano 20K for seven each.
Fifty-four boards, a hundred and twelve configurations, five toolchains. The zoo is not a problem to be eliminated — it is the reality of teaching digital design, and the scripts exist to hide it.
All this matters because of how the project is actually used. When I run a seminar, the room does not contain one kind of board. It contains whatever the university owns, whatever the students brought, and whatever I carried in with me. A student with a Tang Nano 9K and a student with a Nexys A7 should be able to run the same lab from the same repository in the same two commands. That is the whole design goal, and every hour spent on the Bash layer buys back a hundred hours of other people not fighting their tools.
Where this has been used, among other places:
- 2022 — American University of Central Asia, Bishkek, Kyrgyzstan
- 2023 — LaLambda, Tbilisi, Georgia
- 2024 — ADA University, Baku, Azerbaijan; Hacker Dojo, Silicon Valley
- 2025 — Universidad Autónoma de Baja California, Tijuana, Mexico; Russian-Armenian University and the Institute for Informatics and Automation Problems, Yerevan, Armenia
The same examples are used by Школа синтеза цифровых схем, the School of Digital Circuit Synthesis, whose seminars run in 25 universities across Russia and Belarus — which makes it the largest single user of these examples.
Educators use BGM because it removes the vendor complexity from the first week of a course. Nobody learns anything about finite state machines from forty minutes of licence dialogs and directory layouts. And a script that fails for an obscure reason does not cost one raised hand in one room; it costs the same confused hour in every room at once.
The problem AMD handed us
Which brings me to the script in question. For years, Vivado installed itself into a directory that looked like this:
<parent>/Xilinx/Vivado/2023.1
Our setup script looked for exactly that, under $HOME, /opt and /tools. Then, starting with release 2024.2, AMD put the version above the product, and the vendor directory is no longer necessarily called Xilinx:
<parent>/AMD/2026.1/Vivado
So the script stopped finding Vivado. Worse, I later learned it did something more embarrassing than failing. If you had an installation with no version directory at all, the script happily enumerated Vivado’s own internal subdirectories as if they were version numbers, sorted them alphabetically, and picked the last one:
XILINX_VIVADO=.../Xilinx/Vivado/tps PATH=...:.../Xilinx/Vivado/tps/bin
It then printed a confident warning that multiple Vivado versions were installed, listing tps, scripts and lib among them. A student hitting this would get a failure message about vivado not being in the path, several steps later and in the wrong place entirely.
The experiment
I keep two clones of the repository side by side: ~/claude and ~/codex. I gave the same task to Claude in one and Codex in the other, on separate branches, with no knowledge of each other. Then I did three things repeatedly:
- Ask each one to implement or improve the search.
- Ask each one to read the other’s clone and write a report comparing the two implementations, with advantages and disadvantages of each.
- Ask each one to improve its own code based on that comparison.
This is where not being a Bash virtuoso stops mattering. I did not have to invent the right implementation. I had to read two arguments about the same code and decide which one was better — and that is a much easier skill to have, and a much easier one to apply at speed.
Round one: they were not the same
Both agents got the layouts right. The interesting differences were elsewhere, and each had found something the other had missed.
Codex had worried about a case I would not have thought of. Under Cygwin and MSYS — which plenty of students use — the sort on the path can be Microsoft Windows’ own sort.exe, which knows nothing about version sorting. So Codex searched /usr/bin/sort and /bin/sort before trusting sort from the path, and it did not merely check that the -V option was accepted. It fed the candidate two version numbers and checked the answer:
printf '2026.10\tb\n2026.2\ta\n' | LC_ALL=C "$sort_candidate" -t $'\t' -k1,1V
# accept this sort only if 2026.2 really comes out before 2026.10
Claude’s version checked only that sort -V did not error. When I had Claude test its own code against a deliberately broken sort, it picked Vivado 2026.9 over 2026.10 and only printed a warning. That is the kind of bug that produces a confused student and a wasted afternoon.
Claude had found things too. These scripts are sourced into the user’s interactive shell rather than executed, so every variable they fail to declare local stays in your session afterwards. Claude measured it: the original script left seven variables behind, its own first attempt left eleven, and Codex left none. Claude then fixed its own, which is the useful part.
And Claude insisted on one behaviour Codex still does not have. If a directory looks like a Vivado installation but its name does not start with a digit, Codex skips it in silence. Claude names it:
warning: ignoring '.../Xilinx/Vivado/mybuild', because the name
of a version directory is expected to start with a digit.
You can also use XILINX_VIVADO variable to specify a directory ...
For a teaching repository, that message is worth real money. The silent version turns a five-second fix into a forum thread.
Round two: they converged
What happened next is the part I did not expect. Each agent read the other’s critique and took the good ideas. Within a couple of rounds both had moved to shell patterns instead of find, both required a version directory beginning with a digit, both hunted for a sort that genuinely works, both dropped support for versionless installations once I decided that case was out of scope, and both settled on the same message style.
By the final round, Claude ran both implementations through one test harness against identical fake installation trees — including an equal-version tie across two vendor directories, a symlinked version directory, and a sabotaged sort — and reported that in eleven out of eleven situations the two scripts selected the identical directory. Two independent implementations, no shared code, same answer every time.
When two solutions disagree, one of them is teaching you something. When they stop disagreeing, you are probably done.
What was left to compare was no longer behaviour. It was size, diagnostics and taste:
| codex | claude | original | |
|---|---|---|---|
| Lines in the file | 284 | 403 | 255 |
| Functions | 3 | 6 | 1 |
| Variables leaked into your shell | 0 | 1 | 7 |
| Explains a directory it skipped | no | yes | no |
Survives a shadowed sort |
yes | yes | n/a |
| Selection disagreements, of 11 | 0 | 0 | — |
Codex’s version is 119 lines shorter with half the functions and no leaked variables. Claude’s is the only one that tells the student why their installation was ignored. That is a genuine engineering trade-off, and it is the kind of decision I am happy to make myself — which is exactly where I wanted to end up.
The reports
All six reports are in the repository, three from each side, written at successive stages of the competition. They are worth reading in order, because you can watch the gap close:
misc/ai_sessions/20261003_xilinx_support
- claude claude_report_1.html
- claude claude_report_2.html
- claude claude_report_3.html
- codex codex_report_1.html
- codex codex_report_2.html
- codex codex_report_3.html
Each report explains, in plain English and in steps, how a Vivado installation gets chosen — which is documentation the project simply did not have before, and which came out of the competition as a side effect.
What I take away from this
A few things, in the order I think they matter.
Competition surfaces blind spots that review does not
If I had asked one agent to write the script and then asked it to review its own work, I would have got a confident report and the Windows sort.exe bug. The second implementation is what made the first one’s weaknesses visible, because it had made different choices rather than looking harder at the same ones.
Make them test, not just build
The most valuable thing either agent did was build a test harness: stub out the helper functions, create fake installation trees, and run both scripts against the same cases. That is how the 2026.9 versus 2026.10 bug stopped being a theory. “It synthesizes” and “it works” are different claims, and an agent will happily give you the first while sounding like it gave you the second.
Judging is a smaller skill than authoring, and it scales
I could not have written Codex’s sort probe from memory. I had no trouble at all recognising that it was better than Claude’s once both were in front of me with an explanation of each. For a maintainer, that asymmetry is the whole opportunity.
Keep the reports
Committing the comparison reports into misc/ai_sessions/ turned out to be more useful than I expected. Six months from now, when somebody asks why the script probes three different sort binaries, the answer is in the repository instead of in my memory.
The script now finds Vivado under Xilinx, AMD and AMDDesignTools directories, in both the old and the new layout, compares version numbers correctly, and says something useful when it finds nothing. It took two robots arguing with each other and a human deciding who was right.