- History - Early days - when unix was being initially developed, characters were represented with 8 bits since memory was a concern - ASCII was created as a lookup table to display readable characters - ASCII only used the bottom 7 bits, so there was an extra bit that could still be used in an 8-bit character - many different developers created their own “extended character sets” with whatever was convenient (accented letters, emojis etc.) - These extended character sets were not standardized at all and exchanging an email or document across computers would often make it partially or fully unreadable. - at the end of the 90s, there were over 60 versions of the extended ASCII table - all of this became a large issue once the internet was invented and transferring documents became common. - Xerox Character Code Standard (XCCS) - 1980 - 16-bit character encoding - encodes characters required for languages using Latin, Arabic, Hebrew, Greek and Cyrillic scripts; Japanese, Chinese and Korean as well as technical symbols - precursor of unicode - **should find more info about how this system worked** - 1987 - Xerox employee Joe Becker with Apple employees Lee Collins and Mark Davis started working on unicode - Becker published preliminary paper in 1988 dubbing this system as’Unicode’ or ‘universal encoding’ - scheme outlined used 16-bit characters - aimed to only contained modernly used characters, and not obsolete ones - unicode was ready by 1990 after more people joined - 1996 introduced much more capacity for characters including hieroglyphs - Unicode consortium - collection of some of the largest tech companies in the world and helps to maintain the standard - non-profit - Function - Unicode uses variable length encoding, meaning that if a character only needs one byte, thats all it will use - rare characters can use up to 32 bits - each letter is represented by a code point - A code point is just an integer value that is associated with a symbol - A code point is U+ and then a hex number after it - you cant just store a code point in memory, and you need an encoding to do that - encodings - initially each character was just split up into two bytes each - this caused a lot of wasted storage with purely english text though - encodings are ways to convert ideas/symbols into data that computers can read and store - UCS-2/UTF-16 - This stores every character in two bytes no matter what - can represent 65,535 characters, but some like ASCII have lots of 0s and extra data not needed - takes up unneeded space for purely ASCII or ASCII-heavy text - since every character is two bytes, a byte order mark (BOM) is needed to convey the byte order - (small first or large first) - some computers are more efficient with one of these - UCS-2 and UTF-16 are a little different, UCS-2 only allows for 16-bit characters and UTF-16 can allow for 20-bit characters - UCS-2 can only encode things in the basic multilingual plain, so it is mostly obsolete. - with UTF-16, surrogate pairs can be used to encode things not in the BMP - UTF-8 - stores english characters in one byte, and anything else in more than that - since UTF-8 is variable length, each character needs something to tell the computer how many bytes a character takes up - if a byte starts with the bit 0, then the rest of the byte is a 7-bit ASCII value, and the whole character is only 1 byte - if a byte starts with a 1, the computer counts the number of 1s there are at the start before a 0 - if the byte starts with 1110, then there are a total of 3 bytes describing the character (including the first one) - every subsequent bit in the first byte starts to describe the character - each subsequent byte starts with 10 to show that it is not the start byte of a character I’m sure we’ve all seen this character (�) at some point, whether on an out of date website or random text file on your computer. What does it really mean though? the simple answer is that the computer cannot read the character that was in it’s place before, but to fully understand why this is the case, we have to go deeper. Let’s start by talking about some history. When unix was first being developed in the 70s, there needed to be a way to represent characters and words in the memory of computers. To do this, developers used the ASCII character encoding standard. As you might know, computers aren’t able to directly store letters and words in their memory, and have to work in bits and bytes instead. ASCII is a lookup table that associates byte values with characters. Each character is represented by a value from 0-127 which is stored in one byte each. You might be thinking to yourself though, “can’t a byte store 256 different values?” This is true, and basic ASCII only takes up 7 out of the 8 bits in a byte or the first 128 possible numbers, which means that the top half of values were left unassigned in standard ASCII. the problem with only storing 128 characters is that ASCII is very limited. If you wanted to store any special characters, or anything from other languages with different alphabets, you would be out of luck. At first, people tried to solve this problem with the use of extended character sets, which added to ASCII using the unused upper portion of the character set. an extra 128 characters wasn’t enough to store every language though, so dozens of different, incompatible extended character sets were created for different languages, dialects or even operating systems and computers. What this meant is that if text was transferred between two computers that used different extended character sets, any characters used not in the base ASCII set would be interpreted with the new, different extended character set and become unreadable. While this was annoying, it didn’t become a large problem until the creation of the internet when information started to be exchanged between computers much more often. At this point, there were over 60 different common extended character sets in use, and something had to be done to bridge the gaps. In 1987, Xerox employee Joe Becker along with Apple employees Lee Collins and Mark Davis started working on unicode. A year later, Becker published a preliminary paper dubbing the new system they were creating as Unicode meaning “universal encoding.” The new system that they were creating aimed to represent and include all modernly used characters across the world. Unicode doesn’t work in the same way as ASCII, and instead of each character having a direct binary equivalent, each character has a unique code point. These code points are unique integers often represented in hexadecimal and preceded by U+. Each code point references a character, although each character might be shown in slightly different ways depending on the font or rendering engine being used. The maximum number of characters that could fit into unicode is a little over 1.1 million, but there are only currently a few hundred thousand being used. To help organize these characters, Unicode is split into 17 different planes of characters with each one containing different types of characters. The first plane is the basic multilingual plane which contains characters for almost all modern languages, and most writing fits into this plain. The next plane is the supplementary multilingual plane, which includes certain historical scripts such as Egyptian hieroglyphics and cuneiform as well as symbols and emojis. After this, the next two planes are the supplementary ideographic plan and the tertiary ideographic plane, and are primarily used for obscure CJK ideographs mostly used in names. After this, the next ten planes have not been assigned characters and are left open for needs in the future. Plane 14 or the supplementary special-purpose plane only has a small amount of characters that are primarily tags used to help render unicode. The final two planes are both private use planes, and are available to private entities to use and assign as they wish on a case-by-case basis. As mentioned earlier, code points in unicode are represented as hexadecimal integers preceded by U+. Within each plane, 4 hexadecimal values are used to refer to a unique character, giving 16^4 or 65,536 possible characters in each plane. When referencing the basic multilingual plane, only four hex digits are used but when referencing another plane, a fifth or sixth hex digit is added directly after the U+ to show what plane it is in in the range from 1-10. An example of this is the sunglasses emoji which has the code point of U+1F60E. The first 1 means that it resides on the supplementary multilingual plane and then the following four hex digits show where it is within that plane. The problem with code points is that on their own, they cannot be directly stored in a computer. this is where encodings come in, which are processes in which the code points are transformed into a form that a computer can store and understand. One of the simplest encodings is UCS-2, which takes the 4-digit hexadecimal value of a unicode character and stores it in two bytes. While this approach is simple, it has a few problems. The first of these is that since it only uses two bytes, it is unable to encode for anything outside of the basic multilingual plane. While most common text is within this plane, using this encoding would remove the ability to use most symbols or emoji. To solve this problem, the encoding UTF-16 can be used which is very similar to UCS-2 but it is able to encode for characters outside of the basic multilingual plane using surrogate pairs. The problem that both of these encoding share though is that when encoding simple latin characters that make up a majority of text being encoded, they are very inefficient. With every character being forced to take up at least two bytes, any common characters that only need one byte are preceded by unnecessary zeros that waste space and make files larger than they need to be. To solve this problem, the encoding UTF-8 was created. UTF-8 is a little bit different than the first two encodings, since it is able to store characters that only need one byte in only one byte, but is also able to store characters that need more bytes as well. This makes it a variable-length encoding and allows it to access any unicode character in any plane. To understand how this variable-length encoding works, we need to zoom in to the bit level of how a character is stored in UTF-8. Let’s first look at the character A, which is assigned to the code point U+0041. This hexadecimal number can be represented in binary as 1000001, and since it is in the first 128 characters it only needs 7 bits to be represented. Since this is the case, a 0 is added to the front to tell the computer that only one byte is needed to store this character and the following seven bits make up the data. This works for any of the characters in the original 128 character ASCII table, which makes this encoding backwards-compatible with ASCII. For any character beyond this, more than one byte is needed and to indicate this a 1 is used as the first bit instead of a 0. In the first byte representing the character, the same number of ones as the total number of bytes needed are used to tell the computer how many bytes to look for in total for this character. For example, the US cent sign is assigned to the code point U+00A2 which needs 2 bytes to be stored with UTF-8. Because of this, the first byte starts with two ones and then a zero to tell the computer two bytes are needed total, and then the following bits contain the start of the data. Every following byte starts with 10 to tell the computer that it is still part of the same character as the previous byte. If we convert 00A2 into binary, and then split it apart we will get the UTF-8 encoding for the cent symbol. For one more example, we can look at the grinning face emoji, which needs 4 bytes to be stored with UTF-8. This means that the first byte starts with 11110 and the three subsequent bytes each start with 10. The binary form of the code point, U+1F600, is then split up and put into the data portions of each of the bytes giving the UTF-8 encoding of this emoji. Because of many of the advantages that it has like reduced file sizes, backwards compatibility with ASCII and range, UTF-8 is the most commonly used unicode encoding.