Friday, September 25, 2026

Unicode CLDR 49 Beta available for specification review

Building construction emoji imageThe Unicode CLDR 49 Beta is now available for specification review and integration testing. The release is planned for October 21th, 2026, but any feedback on the specification needs to be submitted well in advance of that date. The beta specification is available at Draft LDML Modifications. See also the Migration section of the new release page.


CLDR provides the key building blocks for software to support the world's languages (dates, times, numbers, sort-order, etc.). All major browsers and all modern mobile phones use CLDR for language support. (See Who uses CLDR?)


Via the Survey Tool, contributors supply data for their languages. This data is a factor in determining which languages are supported on mobile phones and computer operating systems.


The beta has already been integrated into the development versions of ICU 79, and ICU4X . We would especially appreciate feedback from non-ICU consumers of CLDR data and on Migration issues. Feedback can be filed at CLDR Requesting Changes.


The following are some of the most significant changes to the CLDR specification (LDML).


  • Clarified the process of selecting the best dateFormatItem when there is no exact match, and how to use appendItems to add missing fields.

  • In the Key/Type Description table, added a description of which keys/types use constructed values and a brief description of the typeValue element.

  • For plural rules, made it clear that they are evaluated in semantic order (zero, then one,…).

  • Part 5: Collation will be updated to accurately reflect the changes that were upstreamed into UTS #10 as part of Unicode 18.0.

  • MessageFormat updates:

    • The :currency and :percent functions are now Stable, with the same implementations as previously.
    • The u:locale option (previously in Draft) has been dropped from the specification.
    • Clarified the computation of the exemplar city for non-location zones.
  • The following new items were added:

    • placeholderBoundarySpacing to allow characters to be inserted where necessary, such as between digits in Chinese dates.
    • Added support for ordinal dates with a new dayOfMonth (and parent elements) for days that are not purely numeric. This is used for ordinal dates, like “Sept 13th, 2026”. (Technical Preview.)
    • Date-Timezone and Time-Day-Of-Week as fallback options for appendItem in dates.
    • intervalFormatRange for constructing ranges (used internally for consistency checks).
    • numericDateSeparator and numericTimeSeparator for customization of numeric dates and times (3-10-2031 → 3/10/2031). (Technical Preview.)
    • alt (alternative) forms of gmtFormat and gmtUnknownFormat to request use of localized equivalents of the term “UTC” instead of “GMT” (limited locales).
    • dualOffsetFormat so that the Localized GMT format can express differences more clearly (e.g., Los Angeles → GMT-8/-7; Phoenix → GMT-7; Denver → GMT-7/-6).
    • numberSystem in currency formats to allow for different formats for locales with multiple number systems.
    • nestedBracketReplacement (48.2) for constructing locale names with parts that have parentheses, e.g., ”birmanês (Mianmar [Birmânia])”.

For a full list of new elements and attributes, see Delta DTDs. There are many more changes that are important to implementations, such as changes to certain identifier syntax and various algorithms. See the Modifications section of the specification for details.


For more details see the draft CLDR 49 release page, which has information on the changes to data and structure, accessing the data, reviewing charts of the changes, and — importantly — Migration issues.


Thursday, September 24, 2026

Recent Unicode Happenings: September 2026

Introducing the Davis Prize and Mark Davis Distinguished Lecture


To read, write, and connect digitally in one’s own language is not a “convenience.” It is a matter of dignity, opportunity, and inclusion.

The Mark Edward Davis Distinguished Lecture and The Davis Prize were created to advance that goal.

Named in honor of Unicode co-founder and Stanford alumnus Dr. Mark Edward Davis—one of the foundational architects of modern multilingual computing—the Lecture and Prize will recognize leaders at any career stage whose work expands digital participation across the world’s languages, scripts, and cultures.

The inaugural Mark Edward Davis Distinguished Lecturer and Davis Prize recipient will be selected in 2027, with the lecture taking place at Stanford University in December 2027.

Nominations open soon. Read the full press release here.

Thank you to Dr. Tom Mullaney and Audrey Gao at Stanford SILICON for their collaboration and leadership in making this recognition for Mark a reality.

Announcing the Unicode® Standard, Version 18.0

The Unicode Standard is the foundation for all digital communications — this major update introduces 13,007 new characters, including three new currency symbols, nine new emoji characters, historical scripts, and many other characters and symbols, bringing the total number of encoded characters to 172,808.

Learn more on the Unicode blog.

Unicode Technology Workshop 2026: Exciting Updates!

Full Program Now Available

Unicode Technology Workshop is coming soon, and this year’s full program is now available online! Upcoming sessions and tutorials will cover a wide range of topics including internationalization, digitally disadvantaged languages, keyboards, emoji, AI, digital humanities, and more. Hear from experts across the internationalization and localization, type design, and language industries worldwide.

If you interact with Unicode technologies, UTW is the event for you. Don't miss your chance to register for one of the world's premier internationalization events — special rates for students, academics, and Unicode members are available.

Registration is flexible, and participants can choose to attend 2 days of tutorials, 2 days of sessions, or all 4 days of UTW.

Tickets are available until Monday, October 12. Don’t miss your chance to register!

Scholarship Applications Now Open Through September 25

We are pleased to be able to sponsor a number of student scholarships for UTW 2026.

Partial scholarships are available to all currently-enrolled students at any level of schooling and early career researchers who have completed their degree within the last three years, and will only cover registration costs. Apply now!

About UTW 2026

Unicode Technology Workshop 2026 (UTW 2026) will take place from October 20-23, 2026 in historic and vibrant Nancy, France. This year’s event will be co-hosted by the Unicode Consortium alongside partners in the Missing Scripts Program, a collaboration among University of California, Berkeley’s Script Encoding Initiative (SEI), Atelier National de Recherche Typographique (ANRT), and Institut Designlabor Gutenberg (IDG).

We look forward to welcoming you to Nancy to build the future of internationalization and multilingual digital communication!

🔗 For the latest updates about UTW 2026, visit https://www.unicode.org/events/utw/2026/

We look forward to seeing you there!

Meet Lighthouse: Coming to Your Keyboard in 2027!


Officially welcoming Lighthouse (U+1F6D9) to The Unicode Standard —  just released in Unicode 18.0!

The Unicode Standard is continually evolving, and very character has a story: here’s what Lighthouse’s proposers had to say about this new emoji.

“Commonly symbolizes orientation, wisdom or guidance in difficult times (“a beacon in the dark”). Lighthouses represent navigation through life challenges, symbolizing the importance of making informed decisions and finding direction during uncertainty. They may also metaphorically represent mentors, friends, or inspirational figures who provide guidance and support.”

Lighthouse will be available for adoption in 2027 — in the meantime, adopt one of our other 156,933 available characters to support Unicode’s mission to ensure everyone can communicate in their own languages across all devices.

About the Unicode Consortium

The Unicode Consortium is the premier 501(c)3 non-profit, open source, open standards body for the Internationalization of software and services. It is arguably the most widely deployed software in the world available across 20 billion devices and counting! At its core, Unicode enables people around the world to communicate in any language.

Your contribution may be eligible for a tax deduction. Please consult with a tax advisor for details.

Wednesday, September 16, 2026

Announcing the Unicode® Standard, Version 18.0

Version 18.0 of the Unicode Standard is now available. The Unicode Standard is the foundation for all digital communications, and this new version supports a wider range of text encoding needs. This major update includes new characters and code charts, updated data files, and updated specifications that define many fundamental aspects of text processing.

This version adds 13,007 new characters, including nine new emoji characters as well as many other characters and symbols, bringing the total number of encoded characters to 172,808.

Among the most anticipated new characters are three new currency symbols:

  • U+20C2 RUFIYAA SIGN
  • U+20C3 UAE DIRHAM SIGN
  • U+20C4 OMANI RIAL SIGN

Each of these symbols was authorized for public use by the respective monetary authority over a year ago, but usage has been hindered by lack of a standardized encoding. With the release of Unicode 18.0, vendors are now able to implement support for these symbols.

The largest set of new characters is for the historical Seal (or “Small Seal”) script. This set of 11,328 ideographic characters has important cultural significance in China, dating back to the Qin Dynasty (around 200 BCE). Another new cultural heritage script from China is Jurchen, used in northeastern China during the Jin Dynasty.

See the delta code charts for details on all the new scripts and characters. For additional details regarding new emoji, see Emoji Recently Added, v18.0

No new algorithms have been introduced in this release, but new data files have been added for Seal and other East Asian scripts, along with a new Unicode Standard Annex documenting these new data files: UAX #60, Data for East Asian Scripts.

Conformance language related to variation selectors has been updated, making clearer which uses of variation selectors are or are not conformant to the Unicode Standard. 

Recommendations for implementations were also added to make non-conformant uses of variation selectors visible in text. This is important as research has shown that sequences of invisible variation selector characters can be used to attack modern AI applications.

For complete details on Unicode Version 18.0, see https://www.unicode.org/versions/Unicode18.0.0/. 




Friday, September 4, 2026

Unicode CLDR 49 Alpha available for testing

Building construction emoji imageThe Unicode CLDR 49 Alpha is now available for integration testing.

CLDR provides key building blocks for software to support the world's languages (dates, times, numbers, sort-order, etc.) For example, all major browsers and all modern mobile phones use CLDR for language support. (See Who uses CLDR?)

The alpha has already been integrated into the development versions of ICU and ICU4X. We would especially appreciate feedback from non-ICU consumers of CLDR data on Migration issues. Feedback can be filed using CLDR Tickets.

Some of the most significant changes in this release include the following (for more detail, see the CLDR 49 release note page):
  • Updated for Unicode 18 including annotations for the new emoji, new scripts, etc.
  • Reflects the most recent updates to external standards and data sources, such as the language subtag registry, UN M49 macro regions, ISO 4217 currencies, etc.
  • New formatting options:
    • New date & time formatting options including:
      • Localized patterns for gluing date & timezoneAppend Items — e.g., Sept 3, EST
      • Ordinal days in dates — e.g., Sept 3rd
      • Customizing numeric separators for dates and times in patterns: 3-10-2031 → 3/10/2031
      • UTC timezone display patterns
      • Dual Standard/Daylight format — Sept 3, UTC+3/+2
      • Structure for preventing digit-digit merges — e.g., '2026/1/29 GMT-817时'
      • Additional skeleton-patterns added for flexible and interval date formats
    • Nested bracket replacement — for constructing locale names with parts that have parentheses, e.g., ”birmanês (Mianmar [Birmânia])”
    • Additional localized locale display name options - see keys
    • New and updated plural and ordinal rules
    • New and improved rule-based number formatting (spellout) rules for many new locales
    • New units — poundal, dyne, milliinch (US mil)
  • Many enhancements of the CLDR specification (LDML) will be in the CLDR 49 Beta (September 23rd).

Locale Coverage Levels

Level Count Regional
Variants
Usage
Modern 100 399 Suitable for full UI internationalization
Moderate 12 14 Suitable for “document content” internationalization, eg. in spreadsheet
Basic 73 99 Suitable for locale selection, eg. choice of language on mobile phone

This is the second release where the new CLDR Organization process is in place for DDL languages. As a result, several locales were able to reach higher levels or had substantial contributions:
  • Adyghe, Kabardian: Adyghe Language and Literature Association
  • Laz: Laz Instittue
  • Ligurian: Council for Ligurian Linguistic Heritage
  • Mara: Mara Language Preservation
  • Coptic: St. Shenouda Coptic Society
± New Level Locales
📈 Modern Akan
📈 Moderate Breton, Coptic
📈 Moderate* Romansh, Shan, Tigrinya
📈 Basic Adyghe, Central Kurdish, Colognian, Kabardian, Kikuyu, Kʼicheʼ, Ladin, Prussian, Qʼeqchiʼ, Sunwar (Sunuwar)
📉 Basic* Tajik, Bashkir, Interlingua, Sardinian, Faroese, Venetian

Note: Each release, the number of items needed for Modern and Moderate increases. So locales without active contributors may drop down in coverage level.

For the details, see the CLDR 49 release note page, which has information on accessing the data, reviewing charts of the changes, and — importantly — will cover Migration issues.