UTF-8 Validation

Given an integer array data representing the data, return whether it is a valid UTF-8 encoding (i.e. it translates to a sequence of valid UTF-8 encoded characters).

A character in UTF-8 can be from 1 to 4 bytes long, subject to the following rules:

  • For a 1-byte character, the first bit is a 0, followed by its Unicode code.
  • For an n-byte character, the first n bits are all 1, the n + 1 bit is 0, followed by n - 1 bytes with the most significant 2 bits being 10.

This is how the UTF-8 encoding works:

Number of Bytes | UTF-8 Octet Sequence (binary)

  • 1 byte: 0xxxxxxx
  • 2 bytes: 110xxxxx 10xxxxxx
  • 3 bytes: 1110xxxx 10xxxxxx 10xxxxxx
  • 4 bytes: 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx

x denotes a bit in the binary form of a byte that may be either 0 or 1.

Note: The input is an array of integers. Only the least significant 8 bits of each integer is used to store the data. This means each integer represents only 1 byte of data.

Example 1
Inputdata = [197,130,1]
Outputtrue
data represents the octet sequence 11000101 10000010 00000001, which is a valid UTF-8 encoding for a 2-byte character followed by a 1-byte character.
Example 2
Inputdata = [235,140,4]
Outputfalse
data represents a 3-byte character start followed by one valid continuation byte, but the second continuation byte does not start with 10, so it is invalid.

Constraints

  • 1 <= data.length <= 2 * 10^4
  • 0 <= data[i] <= 255

Asked at 4 companies

</>

Your Solution

(Ctrl/Cmd + Enter)

Switching Language

Loading template...

Loading...

Sign in to save your progress

AI code evaluation

Get a correctness verdict, missed edge cases, and complexity analysis of your solution.

Sign in to evaluate