Skip to content

docs(blog): add benchmark part 2 post on Plinth workflow hypotheses - #1174

Merged
jabrena merged 3 commits into
mainfrom
feature/benchmark-p2
Sep 3, 2026
Merged

docs(blog): add benchmark part 2 post on Plinth workflow hypotheses#1174
jabrena merged 3 commits into
mainfrom
feature/benchmark-p2

Conversation

@jabrena

@jabrena jabrena commented Sep 2, 2026

Copy link
Copy Markdown
Owner

Summary

Adds the second article in the benchmark series that puts a number on the value of Plinth's AI-native development workflow for Java.

  • Introduces a scenario5 A/B that runs the same OpenSpec technical plan as scenario4 implemented directly, isolating the /implement-spec command from the plan itself
  • Adds a second problem (persistence-and-scheduling "Greek Gods API") so every finding can be read per problem
  • Re-tests all three hypotheses on the expanded corpus (54 → 217 non-template records)

Changes

  • New source: site-generator/content/blog/2026/09/validating-hypotheses-about-plinth-workflow-with-a-benchmark-part-2.md
  • Regenerated docs/ site output via ./mvnw clean generate-resources -pl site-generator -P site-update (index, pagination, archive, feed, sitemap, tag pages, new post page)

Validation

  • ./mvnw clean generate-resources -pl site-generator -P site-update — docs/ diff reviewed, all changes trace to the new post

🤖 Generated with Claude Code

Add the second article in the benchmark series, isolating the
/implement-spec command from the OpenSpec plan via a new scenario5
A/B, adding a second problem, and re-testing all three hypotheses on
the expanded 217-record corpus. Regenerate docs/ site output.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@jabrena jabrena mentioned this pull request Sep 2, 2026
jabrena and others added 2 commits September 3, 2026 10:53
Add scenario/problem overview tables, per-problem skill frequency
breakdowns for S4 vs S5, and trim the recommendations. Regenerate
docs/ output from the updated source.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BeUtLFZcUoP66qo4XmZkGu
Break the S4/S5 testing-skill usage down by tool (claude-code, codex,
cursor), note the Plinth v0.18.0 run, split the POM file-count row, and
tighten prose. Regenerate docs/ output.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BeUtLFZcUoP66qo4XmZkGu
@jabrena
jabrena merged commit 0282cc5 into main Sep 3, 2026
5 checks passed
@jabrena
jabrena deleted the feature/benchmark-p2 branch September 3, 2026 15:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant