UTF-8 Validation
Given an integer array data representing the data, return whether it is a valid UTF-8 encoding (i.e. it translates to a sequence of valid UTF-8 encoded characters).
A character in UTF-8 can be from 1 to 4 bytes long, subject to the following rules:
- For a 1-byte character, the first bit is a
0, followed by its Unicode code. - For an n-byte character, the first
nbits are all1, then + 1bit is0, followed byn - 1bytes with the most significant2bits being10.
This is how the UTF-8 encoding works:
Number of Bytes | UTF-8 Octet Sequence (binary)
1byte:0xxxxxxx2bytes:110xxxxx 10xxxxxx3bytes:1110xxxx 10xxxxxx 10xxxxxx4bytes:11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
x denotes a bit in the binary form of a byte that may be either 0 or 1.
Note: The input is an array of integers. Only the least significant 8 bits of each integer is used to store the data. This means each integer represents only 1 byte of data.
Example 1
Input
data = [197,130,1]Output
truedata represents the octet sequence 11000101 10000010 00000001, which is a valid UTF-8 encoding for a 2-byte character followed by a 1-byte character.Example 2
Input
data = [235,140,4]Output
falsedata represents a 3-byte character start followed by one valid continuation byte, but the second continuation byte does not start with 10, so it is invalid.Constraints
- 1 <= data.length <= 2 * 10^4
- 0 <= data[i] <= 255