v8/unittests at be3c2cdd8de464dd0832c0ba4c9159ce5a0ce979 - v8

History

Justin Ridgewell cedec225c9 Implement DFA Unicode Decoder This is a separation of the DFA Unicode Decoder from https://chromium-review.googlesource.com/c/v8/v8/+/789560 I attempted to make the DFA's table a bit more explicit in this CL. Still, the linter prevents me from letting me present the array as a "table" in source code. For a better representation, please refer to https://docs.google.com/spreadsheets/d/1L9STtkmWs-A7HdK5ZmZ-wPZ_VBjQ3-Jj_xN9c6_hLKA - - - - - Now for a big copy-paste from 789560: Essentially, reworks a standard FSM (imagine an array of structs) and flattens it out into a single-dimension array. Using Table 3-7 of the Unicode 10.0.0 standard (page 126 of http://www.unicode.org/versions/Unicode10.0.0/ch03.pdf), we can nicely map all bytes into one of 12 character classes: 00. 0x00-0x7F 01. 0x80-0x8F (split from general continuation because this range is not valid after a 0xF0 leading byte) 02. 0x90-0x9F (split from general continuation because this range is not valid after a 0xE0 nor a 0xF4 leading byte) 03. 0xA0-0xBF (the rest of the continuation range) 04. 0xC0-0xC1, 0xF5-0xFF (the joined range of invalid bytes, notice this includes 255 which we use as a known bad byte during hex-to-int decoding) 05. 0xC2-0xDF (leading bytes which require any continuation byte afterwards) 06. 0xE0 (leading byte which requires a 0xA0-0xBF afterwards then any continuation byte after that) 07. 0xE1-0xEC, 0xEE-0xEF (leading bytes which requires any continuation afterwards then any continuation byte after that) 08. 0xED (leading byte which requires a 0x80-0x9F afterwards then any continuation byte after that) 09. 0xF1-F3 (leading bytes which requires any continuation byte afterwards then any continuation byte then any continuation byte) 10. 0xF0 (leading bytes which requires a 0x90-0xBF afterwards then any continuation byte then any continuation byte) 11. 0xF4 (leading bytes which requires a 0x80-0x8F afterwards then any continuation byte then any continuation byte) Note that 0xF0 and 0xF1-0xF3 were swapped so that fewer bytes were needed to represent the transition state ("9, 10, 10, 10" vs. "10, 9, 9, 9"). Using these 12 classes as "transitions", we can map from one state to the next. Each state is defined as some multiple of 12, so that we're always starting at the 0th column of each row of the FSM. From each state, we add the transition and get a index of the new row the FSM is entering. If at any point we encounter a bad byte, the state + bad-byte-transition is guaranteed to map us into the first row of the FSM (which contains no valid exiting transitions). The key differences from Björn's original (or his self-modified) DFA is the "bad" state is now mapped to 0 (or the first row of the FSM) instead of 12 (the second row). This saves ~50 bytes when gzipping, and also speeds up determining if a string is properly encoded (see his sample code at http://bjoern.hoehrmann.de/utf-8/decoder/dfa/#performance). Finally, I've replace his ternary check with an array access, to make the algorithm branchless. This places a requirement on the caller to 0 out the code point between successful decodings, which it could always have done because it's already branching. R=marja@google.com Bug: Change-Id: I574f208a84dc5d06caba17127b0d41f7ce1a3395 Reviewed-on: https://chromium-review.googlesource.com/805357 Commit-Queue: Justin Ridgewell <jridgewell@google.com> Reviewed-by: Marja Hölttä <marja@chromium.org> Reviewed-by: Mathias Bynens <mathias@chromium.org> Cr-Commit-Position: refs/heads/master@{#50012}		2017-12-11 21:36:13 +00:00
..
api	Enable clang's -Wunreachable-code warning.	2017-12-04 13:09:25 +00:00
asmjs	Normalize casing of hexadecimal digits	2017-12-02 01:24:40 +00:00
base	Normalize casing of hexadecimal digits	2017-12-02 01:24:40 +00:00
compiler	Enable clang's -Wunreachable-code warning.	2017-12-04 13:09:25 +00:00
compiler-dispatcher	[test] Add TaskRunners to the platform in the compiler dispatcher tests	2017-11-20 15:54:11 +00:00
heap	[heap] Add background GC tracing infrastructure.	2017-12-04 17:28:41 +00:00
interpreter	Normalize casing of hexadecimal digits	2017-12-02 01:24:40 +00:00
libplatform	Reland "[platform] Implement TaskRunners in the DefaultPlatform"	2017-11-15 12:35:54 +00:00
parser	Normalize casing of hexadecimal digits	2017-12-02 01:24:40 +00:00
wasm	[wasm] s/wasm-heap/wasm-code-manager	2017-12-05 16:30:06 +00:00
zone	[heap] Simplify and linearly scale ResourceConstraints::ConfigureDefaults.	2017-05-23 17:00:57 +00:00
bigint-unittest.cc	Normalize casing of hexadecimal digits	2017-12-02 01:24:40 +00:00
BUILD.gn	[wasm] s/wasm-heap/wasm-code-manager	2017-12-05 16:30:06 +00:00
cancelable-tasks-unittest.cc	Make CancelableTask ids unique	2017-08-02 16:10:42 +00:00
char-predicates-unittest.cc	Use ICU for ID_START, ID_CONTINUE and WhiteSpace check	2017-06-14 20:32:49 +00:00
code-stub-assembler-unittest.cc	Remove ComputeFlags, simply pass in Code::Kind instead of Code::Flags	2017-09-29 15:37:27 +00:00
code-stub-assembler-unittest.h	[csa] Add constant folding more universally to CodeAssembler operators	2017-09-12 10:03:10 +00:00
counters-unittest.cc	[runtime] Use methods instead of static functions in RuntimeCallStats.	2017-11-30 12:39:39 +00:00
DEPS	Move unit tests to test/unittests.	2014-10-01 08:34:25 +00:00
detachable-vector-unittest.cc	[cleanup] Replace List with std::vector in api.	2017-09-28 09:32:18 +00:00
eh-frame-iterator-unittest.cc	Normalize casing of hexadecimal digits	2017-12-02 01:24:40 +00:00
eh-frame-writer-unittest.cc	Normalize casing of hexadecimal digits	2017-12-02 01:24:40 +00:00
locked-queue-unittest.cc	Add lock-based unbounded queue	2015-11-18 10:54:13 +00:00
object-unittest.cc	[runtime] Extend InstanceType to uint16_t range of values.	2017-11-22 19:14:09 +00:00
register-configuration-unittest.cc	[Turbofan] Add concept of FP register aliasing on ARM 32.	2016-10-26 16:04:33 +00:00
run-all-unittests.cc	[cleanup] use unique_ptr for the DefaultPlatform	2017-11-14 09:57:18 +00:00
source-position-table-unittest.cc	Decouple SourcePositionTableBuilder from Zone	2017-11-21 12:56:13 +00:00
test-helpers.cc	[unittests] Add TestWithIsolate::RunJS helper method	2017-11-13 14:27:51 +00:00
test-helpers.h	[unittests] Add TestWithIsolate::RunJS helper method	2017-11-13 14:27:51 +00:00
test-utils.cc	[RCS] Add explicit tests for function callbacks	2017-11-14 09:48:08 +00:00
test-utils.h	[RCS] Add explicit tests for function callbacks	2017-11-14 09:48:08 +00:00
testcfg.py	[test] Move GoogleTestSuite to a separate testcfg.	2017-12-11 20:57:07 +00:00
unicode-unittest.cc	Implement DFA Unicode Decoder	2017-12-11 21:36:13 +00:00
unittests.gyp	[wasm] s/wasm-heap/wasm-code-manager	2017-12-05 16:30:06 +00:00
unittests.isolate	[test] Move GoogleTestSuite to a separate testcfg.	2017-12-11 20:57:07 +00:00
unittests.status	Enable RCS unittests again	2017-11-10 09:40:23 +00:00
utils-unittest.cc	Reland "MIPS[64] Implementation of MSA instructions on builtin simulator"	2017-11-28 13:43:23 +00:00
value-serializer-unittest.cc	Enable clang's -Wunreachable-code warning.	2017-12-04 13:09:25 +00:00