What's new in Unicode 7.0?(babelstone.blogspot.co.uk)
babelstone.blogspot.co.uk
What's new in Unicode 7.0?
http://babelstone.blogspot.co.uk/2013/10/whats-new-in-unicode-70.html
6 comments
The same scenario - S and T with cedillas, used in Turkish and (used to be) used in Romanian. Only that Romanian letters should be with comma, not cedillas (on small scale the difference is not distinguishable). So it's Romanian vs Turkish? No, it isn't! S and T with commas were added for Romanian later (somewhere waaay separate from other Romanian diacritics) in the Unicode set and everybody is happy! ...and sane.
Latvian and Marshallese is slightly more confusing, because Latvian commas are already called cedillas in the unicode names, despite not actually being cedillas.
Sure, they'll just extend by adding something like "MARSHALLESE CEDILLA" to refer to actual cedillas. But it still must be frustrating to discover areas where traditional orthographic names are ambiguous or even misleading.
Sure, they'll just extend by adding something like "MARSHALLESE CEDILLA" to refer to actual cedillas. But it still must be frustrating to discover areas where traditional orthographic names are ambiguous or even misleading.
Isn't Linear A a little bit of an overkill? I mean, if we at least knew what it meant, maybe then, but now?
Presumably it's useful to people who are writing papers in the process of deciphering it. Assuming anyone other than crackpots are still trying.
And "MAN IN BUSINESS SUIT LEVITATING" is not overkill?
Seems to devalue the entire standard and project.
edit: Unicode should be universally useful. This is not. I would love to be convinced otherwise however .....
Seems to devalue the entire standard and project.
edit: Unicode should be universally useful. This is not. I would love to be convinced otherwise however .....
You want each character, on it's own, to be universally useful? That's a pretty high standard, there wouldn't be very many characters to meet it and include in the repertoire.
Pretty much every character that exists is useful to some people some time, but not every person all the time.
Pretty much every character that exists is useful to some people some time, but not every person all the time.
... and emoji, in particular, are actually very useful (and well-used) for average every-day communication amongst huge numbers of people.
Obviously some are going to be more useful than others, but it would be quite difficult to find a subset of them that are obviously so completely useless that they should be dropped.
The Unicode consortium's practice of just adopting complete scripts wholesale is vastly simpler and works out quite well in practice.
Obviously some are going to be more useful than others, but it would be quite difficult to find a subset of them that are obviously so completely useless that they should be dropped.
The Unicode consortium's practice of just adopting complete scripts wholesale is vastly simpler and works out quite well in practice.
That's not what I meant. Universal was a too strong, agreed. But the sentiment I meant to express was that codepoints should not be whimsical, tied to any one particular proprietary company or sub-group.
Even if this specific character isn't of any particular use, the fact that you can now represent every character in Webdings in Unicode is kind of nice, I guess.
What's so special about Webdings? I would have thought that Unicode be used for representing the common natural language alphabetic, ideographic and symbolic glyphs that had widespread communal and historical use. Not for encoding some proprietary font by some corporation, that's not how I see it. I'd ask not to be voted down if you disagree, just explain to me why you think otherwise.
"...symbolic glyphs that had widespread communal and historical use"
The answer lies within your question.
The answer lies within your question.
I do not have a problem with generic dingbats that have a genuine historical authenticity. I would have a problem with something designed less than 20 years ago that seems to have very arbitrary and whimsical glyphs as is the case with Webdings which came into being in 1997. It is even more troublesome and annoying that it is the product of some corporation and not a shared cultural artifact.
[deleted]
From today I will use "MAN IN BUSINESS SUIT LEVITATING" as the representation of sending a patch upstream.
🕴
🕴
To be fair, I think it's supposed to be a stylized exclamation mark.
Unicode should be easy to standardize. It's never going to be, but making arbitrary decisions on what's 'useful' or not just makes the process harder by making everyone's blood pressure go up every single time that kind of decision has to be made.
Frankly, one of the biggest advantages of having so many codepoints is being able to burn some of them on less-useful shit simply so the standards folk can move on.
Frankly, one of the biggest advantages of having so many codepoints is being able to burn some of them on less-useful shit simply so the standards folk can move on.
I definitely see what you are saying. I suppose after you've included the glyphs for the world's natural languages and common shared symbols you're going to move into more contentious areas. I'd prefer them to be a bit more conservative though but I see your point of view. What if we start having well known brand logos like the ubiquitous golden arches as codepoints. I would argue they've gone too far here and broken with the spirit of the enterprise.
> What if we start having well known brand logos like the ubiquitous golden arches as codepoints.
I think trademark laws would stop that. I'm not sure, but I don't think anyone else is, either, so the uncertainty would prevent it from happening.
I do see what you're saying, and it would be a slap in the face if a script for a fictional language were standardized before every script for every living language made it in to the standard. However, and this is where you and I differ, I don't really see that much difference between emoticons and other little-used typographical marks living languages sometimes use. Face it: The smiley face is typed a lot more often than the per mil mark (like the percent mark, except 'per thousand' instead of 'per hundred'):
http://en.wikipedia.org/wiki/Per_mil
So the middle finger is rude. People are rude sometimes, and the middle finger is a well-established way of showing that, myths about its origin aside:
http://www.snopes.com/language/apocryph/pluckyew.asp
I think trademark laws would stop that. I'm not sure, but I don't think anyone else is, either, so the uncertainty would prevent it from happening.
I do see what you're saying, and it would be a slap in the face if a script for a fictional language were standardized before every script for every living language made it in to the standard. However, and this is where you and I differ, I don't really see that much difference between emoticons and other little-used typographical marks living languages sometimes use. Face it: The smiley face is typed a lot more often than the per mil mark (like the percent mark, except 'per thousand' instead of 'per hundred'):
http://en.wikipedia.org/wiki/Per_mil
So the middle finger is rude. People are rude sometimes, and the middle finger is a well-established way of showing that, myths about its origin aside:
http://www.snopes.com/language/apocryph/pluckyew.asp
In the future, Man in Business Suit Levitating will be the only form of communication.
Linear A at least has many artifacts.
In Unicode 5.1, a script with a single artifact had the character set from its short inscription included:
http://en.wikipedia.org/wiki/Phaistos_Disc#Unicode
Perhaps your font supports it:
𐇑𐇛𐇜𐇐𐇡𐇽 | 𐇧𐇷𐇛 | 𐇬𐇼𐇖𐇽 | 𐇬𐇬𐇱 | 𐇑𐇛𐇓𐇷𐇰 | 𐇪𐇼𐇖𐇛 | 𐇪𐇻𐇗 | 𐇑𐇛𐇕𐇡[.] | 𐇮𐇩𐇲 | 𐇑𐇛𐇸𐇢𐇲 | 𐇐𐇸𐇷𐇖 | 𐇑𐇛𐇯𐇦𐇵𐇽 | 𐇶𐇚 | 𐇑𐇪𐇨𐇙𐇦𐇡 | 𐇫𐇐𐇽 | 𐇑𐇛𐇮𐇩𐇽 | 𐇑𐇛𐇪𐇪𐇲𐇴𐇤 | 𐇰𐇦 | 𐇑𐇛𐇮𐇩𐇽 | 𐇑𐇪𐇨𐇙𐇦𐇡 | 𐇫𐇐𐇽 | 𐇑𐇛𐇮𐇩𐇽 | 𐇑𐇛𐇪𐇝𐇯𐇡𐇪 | 𐇕𐇡𐇠𐇢 | 𐇮𐇩𐇛 | 𐇑𐇛𐇜𐇐 | 𐇦𐇢𐇲𐇽 | 𐇙𐇒𐇵 | 𐇑𐇛𐇪𐇪𐇲𐇴𐇤 | 𐇜𐇐 | 𐇙𐇒𐇵 |
In Unicode 5.1, a script with a single artifact had the character set from its short inscription included:
http://en.wikipedia.org/wiki/Phaistos_Disc#Unicode
Perhaps your font supports it:
𐇑𐇛𐇜𐇐𐇡𐇽 | 𐇧𐇷𐇛 | 𐇬𐇼𐇖𐇽 | 𐇬𐇬𐇱 | 𐇑𐇛𐇓𐇷𐇰 | 𐇪𐇼𐇖𐇛 | 𐇪𐇻𐇗 | 𐇑𐇛𐇕𐇡[.] | 𐇮𐇩𐇲 | 𐇑𐇛𐇸𐇢𐇲 | 𐇐𐇸𐇷𐇖 | 𐇑𐇛𐇯𐇦𐇵𐇽 | 𐇶𐇚 | 𐇑𐇪𐇨𐇙𐇦𐇡 | 𐇫𐇐𐇽 | 𐇑𐇛𐇮𐇩𐇽 | 𐇑𐇛𐇪𐇪𐇲𐇴𐇤 | 𐇰𐇦 | 𐇑𐇛𐇮𐇩𐇽 | 𐇑𐇪𐇨𐇙𐇦𐇡 | 𐇫𐇐𐇽 | 𐇑𐇛𐇮𐇩𐇽 | 𐇑𐇛𐇪𐇝𐇯𐇡𐇪 | 𐇕𐇡𐇠𐇢 | 𐇮𐇩𐇛 | 𐇑𐇛𐇜𐇐 | 𐇦𐇢𐇲𐇽 | 𐇙𐇒𐇵 | 𐇑𐇛𐇪𐇪𐇲𐇴𐇤 | 𐇜𐇐 | 𐇙𐇒𐇵 |
Doesn't it open up some new automated translation opportunities? Even just by making them easier, if not actually making them possible for the first time.
How would it do that? Just use a custom (private) code block, and distribute a font people can install that will enable it. It's not necessary to make it global.
There is a lot of Linear A text known, and some of it has been decoded. It also shares symbols and even a few words with Linear B. Many books and papers have been written about it. That makes it a reasonable candidate for Unicode inclusion. If nothing else Unicodization will let scholars use standard editors.
I see one character that will be hugely successful now that it's standardized: 1f595
The best part is that since it's not in any fonts yet, if you use it in a document it's essentially a time bomb (or a time-f-bomb if you will).
Surprising that it took this long to introduce the middle finger as a standard character.
Emails will never be the same again.
Emails will never be the same again.
....................../´¯/)
....................,/¯../
.................../..../
............./´¯/'...'/´¯¯`·¸
........../'/.../..../......./¨¯\
........('(...´...´.... ¯~/'...')
.........\.................'...../
..........''...\.......... _.·´
............\..............(
..............\.............\...
....................,/¯../
.................../..../
............./´¯/'...'/´¯¯`·¸
........../'/.../..../......./¨¯\
........('(...´...´.... ¯~/'...')
.........\.................'...../
..........''...\.......... _.·´
............\..............(
..............\.............\...
Still no Klingon? Quvatlh!
Here's an interesting response[1] from Michael Everson back in 1997. Also interesting is that there's already a reserved area in the CSUR (ConScript Unicode Registry) for it[2], and you may even have a font that shows it[3]! I was surprised I had one: !
[1]. http://www.unicode.org/mail-arch/unicode-ml/Archives-Old/UML...
[2]. http://www.evertype.com/standards/csur/klingon.html
[3]. http://www.wazu.jp/gallery/Test_Klingon.html
[1]. http://www.unicode.org/mail-arch/unicode-ml/Archives-Old/UML...
[2]. http://www.evertype.com/standards/csur/klingon.html
[3]. http://www.wazu.jp/gallery/Test_Klingon.html
Also sadly missing the Tengwar of Fëanor, though they're in the Unicode roadmap.
http://www.unicode.org/roadmaps/smp/
There are not enough elves in the Unicode committee.
http://www.unicode.org/roadmaps/smp/
There are not enough elves in the Unicode committee.
At least they introduced the Vulcan salute at U+1F596!
They lost me when they went to 21 bits.
Apparently the Marshallese require some characters to display with cedillas, and Latvian has characters traditionally named "(some letter) WITH CEDILLA" even though they are displayed with commas... so if you just say "LETTER WITH CEDILLA" it's now not clear whether you mean cedilla or comma, and correcting it would break Latvian.
http://www.unicode.org/L2/L2013/13128-latvian-marshal-adhoc....
Not sure how they do this work without going slowly insane.