Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Archival draft — DO NOT MERGE. Keep this PR in draft and unmerged.
Archives the standalone simdutf-swift 1.0.1 String API benchmark harness, recorded samples, readable result tables, and reproduction scripts under
Benchmarks/StringAPIs/.Full benchmark report · Benchmark code · Per-case CSV · Raw samples and metadata
Results
Measured on Apple M2 Pro, macOS 27.0, Swift 6.4, using simdutf's ARM NEON backend. The harness pins the released 1.0.1 dependency independently of the parent package checkout.
Unicode cases use Latin, CJK, emoji, and mixed text at nominal sizes of 1,024–1,048,576 UTF-16 code units. Ranges below are minimum and maximum case-median speedups across those corpora and sizes. ASCII peaks come from separate ASCII cases; the full report identifies their exact sizes and timings. Allocation, buffer borrowing, and output destruction are included.
Default release: Swift
-O, C++-OsOpt-in comparison: Swift
-O, C++-O3Ordinary Swift baselines are
String(decoding:as:)for constructors,Array(string.utf16)for UTF-16 buffers, andstring.unicodeScalars.map { $0.value }for UTF-32 buffers. Across 64 equally weighted valid non-ASCII cases, geometric-mean speedups are 4.24× default and 4.63× with C++-O3.The report also compares temporary-buffer filling, public
transcode, and the newerUTF8Spanscalar iterator. Against the fastest measured standard-library comparator for each case, the overall geometric means are 3.08× and 3.41×.The default 62.74× peak converts a 16 KiB ASCII String into UTF-32, versus the ordinary scalar-map/Array approach: 61.870 µs → 0.986 µs. The two process ratios were 62.42× and 62.21×; the headline uses the ratio of pooled medians.
Tiny outgoing strings can favor standard Swift temporary buffers.
String(decodingUTF8:)uses Swift's decoder and is effectively unchanged. The report separates malformed-input, Foundation, and empty-string measurements. These are warm API microbenchmarks on the stated machine and toolchain.Reproduce
cd Benchmarks/StringAPIs python3 run.py python3 report.pyFresh measurements and their generated report go into ignored
Results/local/; archived samples stay intact. Both optimization modes build before serial timing begins.To regenerate the archived tables without running benchmarks:
Validation
git diff --checkpasses.