Monday, July 23, 2012

Unicode Security Mechanisms, Version 3 Released

Version 3.0 of UTS #39, Unicode Security Mechanisms has been released by the Unicode Consortium, together with a new version of the associated UTR #36, Unicode Security Considerations. Because the Unicode Standard contains such a large number of characters for the writing systems of the world, caution is necessary to avoid exposing programs and systems to possible security attacks. These revised documents describe security considerations for Unicode and specify improved mechanisms for reducing the risk of problems.

Version 3.0 is a major revision. Significant changes include:
  • Mixed Script Detection has extensive revisions to its specification.
  • Restriction Level now has an explicitly defined process.
  • Mixed Number Detection now has an explicitly defined process.
  • Conformance requirements have been extended to include Restriction Level and Mixed Number Detection.
http://www.unicode.org/reports/tr36/
http://www.unicode.org/reports/tr39/

Thursday, July 19, 2012

Version 15 of UTS #18, Unicode Regular Expressions has been released by the Unicode Consortium. Regular expressions are used throughout much of the world's software for matching and manipulating text. UTS #18 provides the foundation for the handling of Unicode text in those expressions.

Version 15 is a major revision. Changes include:
  • Conformance clauses dealing with non 1:1 equivalences were either retracted or modified.
  • A Level 2 conformance clause for full properties was added.
  • New properties, including Name_Alias matching and Script_Extensions, were added.
  • A recommended compact form of Unicode escapes was added: \u{...}.
  • There were many clarifications of the text. See http://www.unicode.org/reports/tr18/tr18-15.html

Monday, July 2, 2012

Dr. Vinton G. Cerf to Keynote IUC 36!

Dr. Vinton G. Cerf, Vice President and Chief Internet Evangelist at Google has just been announced as the keynote speaker for the 36th Internationalization & Unicode Conference. Dr. Cerf has served as vice president and chief Internet evangelist for Google since October 2005. In this role, he is responsible for identifying new enabling technologies to support the development of advanced, Internet-based products and services from Google. Dr. Cerf is widely known as one of the “Fathers of the Internet,” for being a co-designer of the TCP/IP protocols. For details please see the on-line announcement: http://www.unicodeconference.org/e/IUC36-07-02-12.htm

Tuesday, June 26, 2012

Proposed updates for Unicode Collation and IDNA

The proposed update of UTS#10 Unicode Collation Algorithm (UCA) modifies the specification for certain edge cases (overlapping contractions), and tightens the requirements for well-formed collation element tables. The detailed descriptions of parametric tailoring options have been removed, and now refer to the corresponding section in LDML. That section adds new explanations and definitions. There are a number of improvements, including additional examples, and some rearrangement of text. See PRI #223

The data has been updated for the Unicode 6.2 beta review, and the associated CollationAuxiliary.txt file in CollationAuxiliary.zip now includes a description of the implicit fractional weight generation and the context syntax. For more details, see Modifications.

There is also a proposed update of UTS #46 Unicode IDNA Compatibility Processing. The data has been updated for the Unicode 6.2 beta review, with minor changes to the text. See PRI #224

Monday, June 25, 2012

Using the Unicode Glossary

The Unicode glossary is useful for people doing documents, specifications, and general-purpose articles. Each of the glossary entries now has a link on it, and clicking on that link exposes it in the address bar of your browser. This makes it easy to add links directly to the Unicode glossary for terms that may be unfamiliar to readers, such as
http://unicode.org/glossary/#grapheme_cluster or
http://unicode.org/glossary/#code_point.

Wednesday, June 13, 2012

Tutorials Announced for IUC 36

Tutorials Announced for 36th Internationalization and Unicode Conference
Santa Clara, Calif., USA; October 22-24, 2012
Mountain View, CA, USA – June 13, 2012 – The Unicode® Consortium today announced the tutorial sessions for the Thirty-sixth Internationalization and Unicode Conference (IUC). IUC 36 will take place in Santa Clara, Calif., USA at the Hyatt Regency Hotel on October 22-24, 2012, sponsored by Adobe. This is the premier conference on technologies and practices for the creation and management of global and multilingual software applications. For more information about the program, please visit http://www.unicodeconference.org/iuc36-tutorials
The Internationalization and Unicode Conference (IUC) covers the latest in industry standards and best practices for bringing software and Web applications to worldwide markets. This annual event focuses on software and Web globalization, bringing together internationalization experts, tools vendors, software implementers, and business and program managers from around the world.
Tutorial Sessions Include:
  • “An Introduction to Writing Systems & Unicode,” by Richard Ishida, Internationalization Activity Lead, W3C
  • “Unicode – A Grand Tour,” by Michael McKenna, International Product Engineer, Zynga, Inc., and Craig Cummings, Globalization Center of Excellence, Rearden Commerce and UTC Vice Chair, Unicode Consortium
  • “Internationalizing Domain Names in Applications (IDNA),” by Amit Gupta, Member Technical Staff, Adobe Systems
  • “Internationalization, An Introduction (Part I: Character Encoding) (Part II: Enabling),” by Addison Phillips, Globalization Architect, Lab 126
  • “Developing an OpenType Font for Complex Scripts Using Fontforge,” by Pravin Dinkar Satpute, Senior Software Engineer, Red Hat
  • “I18N in Javascript with iLib,” by Edwin Hoogerbeets, Independent Globalization Consultant
  • “Keyboard Design for Tavultesoft Keyman and Unicode,” by Marc Durdin, CEO, Tavultesoft Pty Ltd
  • “Web Internationalization – Standards and Best Practices,” by Tex Texin, Chief Globalization Architect, Rearden Commerce, Inc.
  • “Using ICU Workshop,” by Steven R. Loomis, Software Engineer, IBM
  • “Internationalization and Localization in Ruby and Ruby on Rails,” by Martin J. Dürst, Professor, Aoyama Gakuin University
  • “The Road to World-Class Starts with World-Ready,” by Michael Kuperstein, Localization Engineer and Loïc Dufresne de Virel, Localization Strategist, Intel Corporation
  • “Building Multilingual Websites in Drupal 7 and Joomla 2.5,” by Jim DeLaHunt, Principal, Jim DeLaHunt & Associates
MultiLingual Magazine is the media sponsor. The early-bird registration deadline is September 7, 2012. Sponsorships and exhibit space are available; for more information on sponsoring contact Ken Berk at ken.berk@omg.org, +1-781-444 0404. For exhibiting questions email event_marketing@omg.org. For all other questions email info@unicodeconference.org.
###
About the Unicode Consortium
The Unicode Consortium is a non-profit organization founded to develop, extend and promote use of the Unicode Standard and related globalization standards.
The membership of the consortium represents a broad spectrum of corporations and organizations in the computer and information processing industry. Members are: Adobe Systems, Apple, Google, Government of Bangladesh, Government of India, IBM, Microsoft, Monotype Imaging, Oracle, Rearden Commerce, SAP, The Society for Natural Language Technology Research, The University of California (Berkeley), Yahoo!, plus well over a hundred Associate, Liaison, and Individual members.
For more information, please contact the Unicode Consortium http://www.unicode.org/contacts.html
About the Event Producer
OMG® is the Event Producer for the Internationalization & Unicode Conferences. OMG is an open membership, not-for-profit consortium that produces and maintains computer industry specifications for interoperable enterprise applications. Our specifications include MDA®, UML®, CORBA®, MOF™, XMI® and CWM™. OMG’s specifications are all available for download by everyone without charge.
For more information about OMG, visit us online at http://www.omg.org.

Thursday, June 7, 2012

CLDR 21.0.2: New T Extensions for language/locale identifiers

New T Extension fields and subfields [RFC 6497] are now available for use in BCP47 and Unicode Locale/Language Identifiers. These T extensions provide for the identification of transforms that can be used for tagging content or requesting resources. The new T extension fields and subfields are defined in the following files, as part of the CLDR 21.0.2 release:
For example:
  • "zh-t-i0-pinyin", to indicate Chinese text generated with a pinyin input method
  • "en-t-k0-dvorak", to identify a Dvorak keyboard for English
  • "it-t-k0-osx-extended", to request an extended Mac keyboard for Italian
The private use subfields can be used for private agreements, such as:
  • "ru-t-en-x0-mobile", to indicate a translation from English to Russian for use on a mobile device, or
  • "ja-t-de-t0-und-x0-medical", to identify a machine translation from German to Japanese with a specialized dictionary for medical terms.
Related to this, there is draft keyboard layout data currently slated for CLDR 22.0: see Draft Keyboard Charts.

Wednesday, June 6, 2012

PRI #231: Bidi Parenthesis Algorithm

The Unicode Technical Committee is seeking feedback on a proposal to enhance the Unicode Bidirectional Algorithm (UAX #9) with additional logic--a bidirectional parenthesis algorithm (BPA)--for processing paired punctuation marks such as parentheses. This proposal is intended to produce better bidi-layout results in common text sequences that involve paired punctuation marks. Details of the proposal, with questions for reviewers and a detailed background document are available through the PRI #231 page:
http://www.unicode.org/review/pri231/