fix(composer): preserve UTF-8-safe excerpts - #106
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Implement a UTF-8-safe 80-character cap for the preserved composer excerpt in the shared owner bin/fm-composer-lib.sh. The current cut -c 1-80 runs under LC_ALL=C, truncates at 80 bytes, and can split a UTF-8 code point, making the saved excerpt not reliably resendable. Keep the implementation portable to Bash 3.2 on macOS and Linux with no new dependencies. Add an executable behavior regression covering 79 ASCII characters followed by a multibyte character under LC_ALL=C and under a UTF-8 locale. All touched tests must pass through bin/fm-test-run.sh. Keep this to one fix in one PR and no more than two review rounds.
What Changed
LC_ALL=Cand a UTF-8 locale.Risk Assessment
✅ Low: The bounded shared-helper change preserves complete valid UTF-8 code points under C while retaining the existing excerpt output contract.
Testing
Darwin Bash 3.2.57 reproduced the base split-code-point failure; the target preserved the complete UTF-8 character in both locale classes. The focused runner exited 0 with no gate skips, and the worktree remained clean. No linters or full-suite tests were run per the assigned phase boundary.
Evidence: final focused fm-test-run output
Source: final focused fm-test-run output
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
✅ **Review** - passed
✅ No issues found.
✅ **Test** - passed
✅ No issues found.
bin/fm-test-run.sh tests/fm-composer-lib.test.shfm_composer_state_outputboundary reproduction underLC_ALL=CandLC_ALL=C.UTF-8against base and targetDirect 2-, 3-, and 4-byte UTF-8 boundary checks underLC_ALL=Cgit status --short --untracked-files=all✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.