core, std: support various [char] matchers for Pattern<&OsStr> - #161765
Draft
pacak wants to merge 17 commits into
Draft
core, std: support various [char] matchers for Pattern<&OsStr>#161765pacak wants to merge 17 commits into
pacak wants to merge 17 commits into
Conversation
Right now things are undertested and underspecified. Some of the library code would get in a loop if searcher starts returning empty rejects. And there's no tests for backwards multi byte char matchers. Pull request I'm reviving had a problem implementing that, so making sure it's tested before the actual code lands. Right now it is possible to break both tests (and user code) without breaking anything else in the test suite I think.
Add a Haystack trait describing something that can be searched in and make core::str::Pattern (and related types) generic on that trait. This will allow Pattern to be used for types other than str (most notably OsStr). This somewhat follows the Pattern API 2.0 design. While that design is apparently abandoned (?), it is somewhat helpful when going for patterns on OsStr, so I’m going with it unless someone tells me otherwise. ;) For now leave Pattern, Haystack et al in core::str::pattern. Since they are no longer str-specific, I’ll move them to core::pattern in future commit. This one leaves them in place to make the diff smaller. @pacak: I moved some (or all new) of the `P: Pattern<&'a str> constraints into where clause to keep things narrower: ``` pub fn foo<'a, P: Pattern<&'a str>>(&'a self, pat: P, ...) ... ``` to ``` pub fn replacen<'a, P>(&'a self, pat: P, ...) ... where P: Pattern<&'a str>, ``` Original code had indices in Haystack abstracted as an associated type Cursor. Replaced with usize - Cursor adds noise with not much value. Changed wording in 2-3 places - for example Searcher is generic over a few types so it makes more sense to talk about split points in general with utf8 split points as an example for `&str`.
Pattern is no longer str-specific, so move it from core::str::pattern module to a new core::pattern module. This introduces no changes in behaviour or implementation. Just moves stuff around and adjusts documentation.
Introduce core::pattern::Split and core::pattern::SplitN internal types which can be used to implement iterators splitting haystack into parts. Convert str’s Split-family of iterators to use them. In the future, more haystacks will use those internal types. Co-authored-by: Peter Jaszkowiak <p.jaszkow@gmail.com> @pacak: Fixed some typos, added a few `#[inline]`. Since there's no `H::Cursor` - I had to add `ctx: PhantomData<H>`.
This reverts commit 85cf233ced0d0fe02734c8a83b6d79ccc5432d06. Gone for now, I'll reimplement it later in str_bytes.rs, will confirm with the benchmarks included that the optimization still applies
Introduce core::pattern::EmptyNeedleSearcher internal type which implements logic for matching an empty pattern against a haystack. Convert core::str::pattern::StrSearcher to use it. In future more implementations will take advantage of it. Also adapt and rework TwoWayStrategy into an internal SearchResult trait which abstracts differences between Searcher’s next, next_match and next_rejects methods. It makes it simpler to write a single generic method implementing optimised versions of all those calls. @pacak: - Fixed a few typos. - There's no H::Cursor parameter so code gets a bit simplified. - Added a test to assert how TwoWaySearcher runs with EmptyNeedleSearcher
@pacak: - made more things const fn - there was a (copy-paste?) error in try_finish_byte_sequence so I added a test that checks try_next_code_point(_reverse) with some values, including invalid ones. - reworded a few comments (passive voice, etc) Also different comments: since former is public and later is private due to historical reasons. > This is different than [`next_code_point`] in that it doesn't assume > This is different than `next_code_point_reverse` in that it doesn't assume
This was referenced Aug 25, 2026
This comment has been minimized.
This comment has been minimized.
pacak
force-pushed
the
push-svrvnxkpuqul
branch
from
August 25, 2026 14:58
f82c889 to
768ca4e
Compare
This comment has been minimized.
This comment has been minimized.
Introduce a new core::str_bytes module with types and functions which handle string-like bytes slices. String-like means that they code treats UTF-8 byte sequences as characters within such slices but doesn't assume that the slices are well-formed. A `str` is trivially a bytes sequence that the module can handle but so is OsStr (which is WTF-8 on Windows and unstructured bytes on Unix). Move bunch of code (most notably implementation of the two-way string-matching algorithm) from core::str to core::str_bytes. Note that this likely introduces regression in some of the str function performance (since the new code cannot assume well-formed UTF-8). This is going to be rectified by following commit which will make it again possible for the code to assume bytes format. This is not done in this commit to keep it smaller. @pacak: - Added a few comments - tried to hide internal types from the diagnostic And then there's two different bugs where it would report matched areas as rejected. This broke str::trim_end_matches and who knows what else. Caught it thanks to tests in the previous commit. And one underflow bug on invalid input.
It works right now, but original implementation of the next commit breaks them with none of existing tests catching this regression.
Since core::str_bytes module cannot assume byte slices it deals with are well-formed UTF-8 (or even WTF-8), the code must be defensive and accept invalid sequences. This eliminates optimisations which would be otherwise possible. Introduce a `Flavour` trait which tags `Bytes` type with information about the byte sequence. For example, if a `Bytes` object is created from `&str` it’s tagged with `Utf8` flavour which gives the code freedom to assume data is well-formed UTF-8. This brings back all the optimisations removed in previous commit. @pacak: - removed IS_WTF8 associated constant - unused - fixed a bug related to multibyte reverse matching: `next_code_point_reverse` reads the input via Iterator::next_back, passing `bytes.iter().rev()` reverses it a second time. Not good.
I reverted `ByteNeedle` change earlier, time to add the same functionality back. `ByteSearcherState` is mostly copied from `CharSearcherState`, does a single ascii byte search. It is possible to do the dispatch inside of a CharSearcherState, but that makes it a bit slower.
pattern::find_str 4775.12ns/iter -> 2561.57ns/iter
pattern::rfind_str 5621.05ns/iter -> 2492.68ns/iter
Implement Haystack for &OsStr and Pattern<&OsStr> for &str, char and Predicate. Furthermore, add prefix/suffix matching/stripping and splitting methods to OsStr type to make use of those patterns. Using OsStr as a pattern is *not* implemented. With rust-lang#118485 - I added find and rfind
pacak
force-pushed
the
push-svrvnxkpuqul
branch
from
August 25, 2026 18:39
768ca4e to
3865faa
Compare
This comment has been minimized.
This comment has been minimized.
To work around orphan rules, introduce a wrapper type for predicate
functions to be used as pattern. Specefically, if we want to add
predicat pattern implementation for OsStr type, doing it with a naked
`FnMut` results in compile-time errors:
error[E0210]: type parameter `F` must be covered by another type when it
appears before the first local type (`OsStr`)
impl<'hs, F: FnMut(char) -> bool> core::pattern::Pattern<&'hs OsStr> for F {
^ type parameter `F` must be covered by another type
when it appears before the first local type (`OsStr`)
std: add predicate pattern support to OsStr
Due to technical limitations adding support for predicate as patterns
on OsStr slices must be done via core::pattern::Predicate wrapper type.
This isn’t ideal but for the time being it’s the best option I've came
up with.
The core of the issue (as I understand it) is that FnMut is a foreign
type in std crate where OsStr is defined.
Using predicate as a pattern on OsStr is the final piece which now
allows parsing command line arguments.
Sadly MultiCharEq needs to be public if we want to keep it in the same place as other related code. Can probably make it sealed. Going via `AsRef<[char]>` comes at about 20% performance penalty
pacak
force-pushed
the
push-svrvnxkpuqul
branch
from
August 25, 2026 20:21
3865faa to
ba2841c
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Sadly MultiCharEq needs to be public if we want to keep it in the same
place as other related code. Can probably make it sealed.
Going via
AsRef<[char]>comes at about 20% performance penaltyThis PR is part of a stack containing 16 PRs:
main