Repository navigation
Conversation
Indentation tracks current_indent_len as a byte index into its indents cache, but grew and shrank it by indent_size, a count of characters. Writer::new_with_indent takes a u8 and char::from turns any byte above 0x7F into a two byte character, so the index landed inside a character and the slice in current() panicked, or one level of indent came out half as wide as asked for. Serializer::indent takes a char and passed it on as `indent_char as u8`, which truncated every non-ASCII character to a byte and emitted control characters instead. Route it through Indentation::with_char so the character survives.
|
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #1017 +/- ##
==========================================
- Coverage 57.31% 55.10% -2.22%
==========================================
Files 46 51 +5
Lines 18197 18822 +625
==========================================
- Hits 10429 10371 -58
- Misses 7768 8451 +683
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The bug
Indentation(src/writer.rs:670-737) keepscurrent_indent_lenas a byte index into itsindentscache --
current()slices with it -- butgrowandshrinkmoved it byindent_size, which is a count ofcharacters. The two agree only while the indent character is one byte long.
Writer::new_with_indent(inner, indent_char: u8, indent_size: usize)is public and takes au8, andchar::from(u8)maps every byte above0x7Fto a character that is two bytes in UTF-8. So the mismatch isreachable from safe code with a plain byte argument:
With an even
indent_sizeit does not panic, it just indents half as far as asked:Separately,
Serializer::indent(src/se/mod.rs:819) accepts acharand then throws most of it away:'\u{2003}' as u8is0x03. Serializing withindent('\u{2003}', 2)producesU+0003 is not a legal XML character (XML 1.0
Charexcludes C0 controls other than tab, LF and CR), so theserializer returns
Okwith a document that a conformant parser must reject. Any indent character aboveU+007F -- an em space, an ideographic space -- goes the same way, and
charis the only type the caller canpass.
The fix
Indentation::with_char(char, usize);Indentation::new(u8, _)becomes a thin wrapper over it, soWriter::new_with_indent's public signature is unchanged.level_len()=indent_size * indent_char.len_utf8(), and use it ingrowandshrinksocurrent_indent_lenstays a byte length.additional(), convert the added level the same way. The value passed in isAttributeIndent::WriteConfigured(i.indent_size)(src/writer.rs:533), and the variant's own doc commentat
src/writer.rs:444says "Write specified count of indent characters" -- so the added amount is acharacter count and needs the same conversion, applied to the added level only, not to the running total.
Serializer::indentatwith_charso thecharit already accepts survives.ASCII is untouched:
len_utf8()is 1, and every arithmetic site reduces to what it was.I kept
new_with_indent'su8rather than widening it tochar, since that would be a breaking change forevery caller. Happy to switch it if you would rather fix the signature at the same time.
Verification
New module
multi_byte_indent_charintests/writer-indentation.rs: nested elements, attribute indent atdepth 0 and at depth 1, an odd
indent_size(the case that panicked), an ASCII case pinning that nothingmoved, and the serializer case.
cargo test --all-featurespasses; so do--no-default-features,--features serialize,--features serialize,encodingand--features serialize,escape-html, andcargo test --all-features --benches --tests.cargo fmt -- --checkclean,cargo +1.86.0 check(the MSRV the workflow pins) clean.boundary is pinned from both sides:
Indentation::new(indent_char as u8, ..)inSerializer::indentfails onlyserializer_indent_char_is_not_truncated;additional()to(current_indent_len + additional_indent) * len_utf8()fails onlyin_attributes_nested.That second one is why
in_attributes_nestedexists. The depth-0 attribute test cannot see the difference --current_indent_lenis 0 there, so both forms give the same answer. It only shows up one level in.AI disclosure
Written with AI assistance (Claude Code, claude-opus-5). Per the AI policy: I have reviewed the change, I can
explain it without going back to the agent, and I will answer review questions in my own words rather than
pasting model output.