Dan Luu tested whether simple instructions on testing techniques and libraries improve the correctness of a Zstd implementation written by coding agents. All implementations were in Rust.
The author tested 26 instruction variants and four skills, ready-made sets of instructions for agents. Fuzzing tests a program with many random inputs.
Property-based testing tests general program properties with different inputs. At xhigh, fuzzing and property-based testing were on average a little better than formal methods.
The mode with no extra instructions performed above average. Testing-related skills recommended by Codex performed worse.
Claim check:
- Dan Luu tested whether simple instructions on testing techniques and libraries improve the correctness of a Zstd implementation written by coding agents. (confirmed by the publication itself: evidence; «Here, we test if simple instructions to agents to use particular techniques or libraries improve implementation correctness.»)
- All implementations were in Rust. (confirmed by the publication itself: evidence; «All implementations were in Rust.»)
- The author tested 26 instruction variants. (confirmed by the publication itself: evidence; «The 26 prompt conditions tested were ACL2, Adaptive (agents asked to use the best technique), Alloy»)
- The author also tested four skills. (confirmed by the publication itself: evidence; «Additional, 4 skills were tested:»)
- At xhigh, fuzzing and property-based testing were on average a little better than formal methods. (confirmed by the publication itself: evidence; «Looking at xhigh, on average, the fuzzing and PBT-related conditions did a little better than formal methods on average»)
- The mode with no extra instructions performed above average. (confirmed by the publication itself: evidence; «However, Default (no additional instructions) does well above average.»)
- Testing-related skills recommended by Codex performed worse. (confirmed by the publication itself: evidence; «The testing-related skills codex recommended we try underperformed»)
Primary sources:
score 68.0 · kind research · revision 1 · stories st-1r8n29x