emoji in unicode is the funniest thing I ever watched happen 

most people who are into writing systems are not into computers and vice-versa. so Unicode drama was a very particular kind of hobby drama. most of you have no idea how funny the current situation is.

see, it gets pretty tricky to define what's a "character" across every writing system in the world. for 20 years, Unicode had very strict rules about what gets into the standard: no "graphical variations", only "abstract characters" are defined. you might think that "A" is a big pointy wedge with an horizontal stroke but from the point of view of Unicode, "A" is a number (65) and a name ("LATIN CAPITAL LETTER A") and some properties ("Letter", "Uppercase", "Narrow", "Left-to-right"). the facts that "A" is usually this triangley thing, but in some styles A looks more like a round oval with a tail, and in blacktext A looks more like an U shape—all that mean nothing to Unicode. all of those are merely graphical variations ("glyphs") of 65, LATIN CAPITAL LETTER A.

this gets of course fuzzy at the edges. The character 骨, "bone", can be written with the top component facing left, or facing right. This is unarguably a "graphic variation"—it's still the same "abstract character"—so Unicode leaves that for the font to decide. Unfortunately it has become custom in the modern era (but not in historical texts) that mainland China draws it to the left, whereas Japan and Taiwan and the others do it to the right. This means that when people changed their systems from legacy encodings to Unicode, suddenly their characters would be drawn in the hated Chinese way / the hated Japanese way, and since Unicode is a "foreign imposition" they would blame foreigner ignorance of their cultural nuances. In fact the Unicode hànzì team was made of specialists from East Asia and the reasoning for this bug was perfectly logical for nerds: if your system is set to Japanese it should be using Japanese fonts with the glyph to the right, it's not our fault if Microsoft/Linux/etc. is getting the wrong fonts. but in your legacy system it would "just work" and look right; upgrade to Unicode, problems happen. the dorks at Unicode would have endless discussions over this and reject the idea of adding "Chinese 骨" and "Japanese 骨" as different characters because that's technically Wrong, these are the same abstract character.

multiply that by a thousand. imagine 20 years of strict gatekeeping like this, with people constantly proposing new characters from this and that source and the guardians of Unicode rejecting them because they can be understood as merely graphical variations of existing characters, or as too obscure to be worthy encoding, etc. (it took so long to get hentaigana into Unicode, like ~8 years between proposal and finally having the characters encoded! you can't represent any premodern manuscripts without it!)

--

then Japanese people start making cellphones with Internet, before everybody else.

and they put "picture" (e) "characters" (moji) because they're cute. each company makes their own emoji set, incompatible on purpose so you have to buy into their special, incompatible little internets.

see, the only situation where Unicode would make exceptions to their strict gatekeeping was "round-trip compatibility". a character like ª is just an especially typeset 'a' and thus should not count as an abstract character. but some previous text encodings like ISO-8859-1 already had an 'ª' separated from 'a'. so if you converted a text file from ISO-8559-1 and Unicode had no 'ª', it would become 'a', and now you can't distinguish it from regular 'a', so there's no way to convert it back to ISO-8859-1. you cannot "round-trip" between encodings. so in this case and in only ☝️ this case, the Consortium would allow breaking the usual gatekeeping rules, carefully and case-by-case.

then a guy from Google asked to add all the Japanese emoji from every vendor. after all, without that you couldn't represent Japanese meeru messages in Unicode and convert them back. the Consortium obviously refused. so Google teamed up with Apple and kinda highkey threw their weight, financially and otherwise, because obviously they wanted this in the nascent smartphone culture, with powerful interests completely overriding the cabal of writing hobbyists who had hitherto kept Unicode so pure, and after years of elaborate technical criteria refusing tons of characters that people proposed for serious use, Unicode 6 (2010) included the first set of emoji.

and now all of those Unix greybeards and linguistics associates just nod sadly in a corner from the purely ceremonial Consortium chairs at whatever Big Tech mandarins came up with this year, like, "yeah sure that's what we need, U+1FAE9: FACE WITH BAGS UNDER EYES 🫩, whatever. U+1FAC8: HAIRY CREATURE 🫈, why not at this point."

Show older
Computer Fairies

Computer Fairies is a Mastodon instance that aims to be as queer, friendly and furry as possible. We welcome all kinds of computer fairies!