🏠 Home • 📖 About • 🚀 Quick Start • 📘 Tutorial • 🌐 Locale • 📋 Changes • 📚 Wiki • 🇹🇷 Türkçe
AmberDB::Locale is a locale-aware text processing engine designed for multilingual applications. Written in Perl, it provides the following core capabilities:
| Capability | Description |
|---|---|
| Case conversion | Locale-aware uc, lc, ucfirst |
| Sorting | Unicode Collation Algorithm (UCA) based |
| ASCII transliteration | Convert accented characters to plain ASCII (slugs, IDs) |
| Number → Text | Written-out numbers for invoices/documents |
| Date/time formatting | Locale-specific date formats |
| Number/currency formatting | Grouping separators, decimal separators, symbol placement |
| HTML entity decoding | Named + numeric entity decode |
| Plural rules | CLDR-based plural form selection |
| UTF-8 safe substring | Character-based slicing (not byte-based) |
Architectural principle: Engine logic lives in AmberDB::Locale.pm, while language data resides in AmberDB::Locale::Lang::* packages as pure data. Engine and data are completely separated.
AmberDB::Locale ← Main engine (all logic here)
├── AmberDB::Locale::Lang::en ← English data (default/fallback)
├── AmberDB::Locale::Lang::tr ← Turkish data
├── AmberDB::Locale::Lang::de ← German data
├── AmberDB::Locale::Lang::fr ← French data
├── AmberDB::Locale::Lang::es ← Spanish data
├── AmberDB::Locale::Lang::ru ← Russian data
├── AmberDB::Locale::Lang::az ← Azerbaijani data
├── AmberDB::Locale::Lang::ar ← Arabic data
└── AmberDB::Locale::Currency ← ISO 4217 universal currency data
Each Lang::* module contains only a single subroutine called data() and returns a hash-ref. No logic is included.
use AmberDB::Locale;
# 1) Named-param API (recommended)
my $lang = AmberDB::Locale->new(language => "tr");
# 2) Hashref API
my $lang = AmberDB::Locale->new({ language => "de" });
# 3) Positional string API
my $lang = AmberDB::Locale->new("fr");
# 4) No args → defaults to "en"
my $lang = AmberDB::Locale->new();
new() call for the same language returns the existing instance (singleton-like)."turkish", "tr_tr", "tr-tr" are automatically mapped to "tr".cluck warning is emitted and it falls back to en.Unicode::Collate::Locale; if it fails, falls back to base Unicode::Collate, and if that’s also unavailable, uses a custom sort_map Schwartzian transform.# Language is automatically pulled from cfg
my $db = AmberDB->new(cfg => { language => "tr" });
$db->uc("ığdır"); # works via inheritance: IĞDIR
data() Hash-Ref)Complete schema of the structure returned by each language module:
{
uc_map => { }, # Character mapping applied before uc()
lc_map => { }, # Character mapping applied before lc()
sort_map => { }, # Custom sort weights (if no collator)
alphabet_chars => "...", # Safe character class for normalize()
accent_map => { }, # normalize() accent → local equivalent
ascii_map => { }, # to_ascii() pre-mapping before NFD
numbers => { }, # num2text number words
html_entities => { }, # Locale-specific extra HTML entities
months => [ ], # 12-element month names
days => [ ], # 7-element day names (Sunday → Saturday)
number_format => { }, # Decimal/grouping separator settings
default_currency => "",
currency_position => "prefix|suffix",
currency_space => 0|1,
date_format => { }, # short, medium, long, full, time, datetime
plural_rule => "...", # CLDR plural rule expression
}
uc($string) — Convert to uppercasemy $tr = AmberDB::Locale->new(language => "tr");
$tr->uc("ığdır"); # "IĞDIR"
$tr->uc("istanbul"); # "İSTANBUL" (i → İ, Turkish rule)
my $de = AmberDB::Locale->new(language => "de");
$de->uc("straße"); # "STRASSE" (ß → SS)
Turkish detail: The mapping
'i' => "\x{130}"insideuc_mapis applied before Perl’sCORE::uc()call. This ensures thei → İconversion is done correctly.
lc($string) — Convert to lowercase$tr->lc("İSTANBUL"); # "istanbul" (İ → i, I → ı)
$tr->lc("IĞDIR"); # "ığdır"
Inside
lc_map, the mappings'I' => "\x{131}"and"\x{130}" => 'i'are applied beforeCORE::lc().
ucfirst($string) — Capitalize word beginnings$tr->ucfirst("istanbul büyükşehir belediyesi");
# "İstanbul Büyükşehir Belediyesi"
First the entire string is lowercased via lc(), then the first character after spaces, periods, exclamation marks, colons, quotes, /, (, ) is uppercased.
fold($string) — Normalization for searchmy $key = $tr->fold("İSTANBUL"); # "istanbul" (NFKC + lc)
Applies Unicode NFKC decomposition + locale lc(). Designed for search indexing and matching.
ieq($str1, $str2) — Case-insensitive comparison$tr->ieq("İstanbul", "istanbul"); # 1 (true)
$tr->ieq("Ankara", "ankara"); # 1 (true)
$tr->ieq("Ankara", "İzmir"); # 0 (false)
If a Unicode::Collate::Locale collator is available, it uses that for comparison; otherwise it compares fold() results with eq.
sort(\@list [, $field])my $tr = AmberDB::Locale->new(language => "tr");
# Simple array sorting
my @sorted = $tr->sort(["İzmir", "Ankara", "Van", "Şanlıurfa", "Bursa", "Çanakkale"]);
# => ("Ankara", "Bursa", "Çanakkale", "İzmir", "Şanlıurfa", "Van")
# Array of hashrefs — sort by field name
my @sorted = $tr->sort(\@products, "name");
# Array of arrayrefs — sort by index number
my @sorted = $tr->sort(\@rows, 2);
Sorting strategy (priority order):
Unicode::Collate::Locale (locale-specific UCA table)Unicode::Collate (base UCA)sort_map Schwartzian transform (last resort)normalize($string) — Clean upmy $clean = $tr->normalize('<p>Kâr & zarar çizelgesi</p>');
# "Kar zarar cizelgesi"
Processing order:
decode_entities)<...> → space)accent_map (e.g., â → a, ô → ö)to_ascii($string [, $nonspace]) — ASCII transliteration$tr->to_ascii("çarşı"); # "carsi"
$tr->to_ascii("İstanbul", 1); # "istanbul" (slug mode)
$tr->to_ascii("Große Straße"); # "Grosse Strasse" (de locale)
my $de = AmberDB::Locale->new(language => "de");
$de->to_ascii("Müller"); # "Mueller" (DIN 5007-2: ü → ue)
Processing order:
lc() (if $nonspace is provided)normalize()ascii_map (characters NFD can’t decompose: ı → i, ß → ss, ä → ae, etc.)\p{M} (combining marks)[a-z0-9,.\-_ ]$nonspace mode: spaces → _, clean underscore repetitionnum2text($number [, %options])my $tr = AmberDB::Locale->new(language => "tr");
$tr->num2text(0); # "Sıfır"
$tr->num2text(1); # "Bir TL"
$tr->num2text(1000); # "Bin TL" (NOT "Bir Bin")
$tr->num2text(100); # "Yüz TL" (NOT "Bir Yüz")
$tr->num2text(1234.56); # "Bin İki Yüz Otuz Dört TL Elli Altı KR"
$tr->num2text(-42); # "Eksi Kırk İki TL"
# With custom currency
$tr->num2text(99.99, currency => { main => "EUR", sub => "cent" });
# With completely custom number data
$tr->num2text(5, numbers => { zero => "Yok", ones => [...], ... });
Supported options:
| Option | Description |
|---|---|
currency => { main => "...", sub => "..." } |
Main/sub currency names |
numbers => \%hash |
Full or partial number word data override |
Locale-specific rules:
| Rule | Description | Example (tr) |
|---|---|---|
hundred_one_prefix |
“One” prefix for 100 | 0 → “Yüz” |
thousand_one_prefix |
“One” prefix for 1000 | 0 → “Bin” |
decimal_sep |
Decimal separator | ',' → “1.234,56” format |
Eastern Arabic digits (
٠١٢٣٤٥٦٧٨٩) and Persian digits (۰۱۲۳۴۵۶۷۸۹) are automatically converted to Western digits.
safe_chars($string)$lang->safe_chars("hello world! @#$ 123");
# "hello world 123"
Strips all foreign or special characters not matching the language’s alphabet_chars definition.
Core morphological and phonetic analysis methods powering AmberDB full-text search engine:
normalize_word($word [, $mode_write])Normalizes a word phonetically and morphologically for search indexing and query matching:
my $tr = AmberDB::Locale->new(language => "tr");
# 1. Query mode (default, mode_write = 0): Strips suffixes / clitics
$tr->normalize_word("Türkiye'nin"); # "turkiye" (apostrophe suffix 'nin' stripped as stop-word)
$tr->normalize_word("Türkiye'de"); # "turkiye"
# 2. Write mode (mode_write = 1): Generates both root and joined compound for indexing
$tr->normalize_word("Türkiye'de", 1); # "turkiye turkiyede"
# 3. Circumflex / accent normalization
$tr->normalize_word("kârın"); # "karin"
$tr->normalize_word("ÂLÎM"); # "alim"
# 4. Word-final consonant devoicing / phonetic assimilation
$tr->normalize_word("tevhid"); # "tevhit" (d$ => t)
$tr->normalize_word("gazab"); # "gazap" (b$ => p)
$tr->normalize_word("mehmed"); # "mehmet" (d$ => t)
search_pattern($query)Converts a search query string into a locale-aware regex pattern matching regional character variants:
my $pattern = $tr->search_pattern("Türkiye");
# Produces regex token pattern matching Turkish and ASCII variants (e.g. "t[uü]rk[iıİI]y[eE]")
search_regex($string, $pattern)Performs a case-insensitive, locale-aware regex match of $pattern inside target $string. Returns 1 on match, 0 otherwise:
my $found = $tr->search_regex("İstanbul Boğazı", "istanbul"); # 1
my $match = $tr->search_regex("İzmir Kordon", $pattern); # 1
decode_entities($string)$lang->decode_entities("& < > € € ç");
# "& < > € € ç"
Three stages:
€ → €€ → €&, <, , €, etc.) + locale-specific extras (ç, š, ğ, etc.)substring($string, [$offset], $length)my $tr = AmberDB::Locale->new(language => "tr");
$tr->substring("Çanakkale", 0, 4); # "Çana" (4 characters, not 4 bytes!)
$tr->substring("İstanbul", 2, 3); # "tan"
# Also safe with raw UTF-8 byte strings:
my $raw = encode('UTF-8', "Şanlıurfa");
$tr->substring($raw, 0, 5); # "Şanlı" (correctly re-encoded)
Checks whether the string is decoded using
Encode::is_utf8(). If it’s raw bytes, thedecode → substr → encodechain is applied; this prevents multibyte characters from being cut in half.
format_date($time [, $pattern_or_style])my $tr = AmberDB::Locale->new(language => "tr");
my $epoch = time(); # e.g.: 2026-08-09
$tr->format_date($epoch); # "09.08.2026"
$tr->format_date($epoch, 'medium'); # "9 Ağu 2026"
$tr->format_date($epoch, 'long'); # "9 Ağustos 2026"
$tr->format_date($epoch, 'full'); # "Pazar, 9 Ağustos 2026"
$tr->format_date($epoch, 'time'); # "14:30"
$tr->format_date($epoch, 'datetime'); # "09.08.2026 14:30"
# Custom pattern
$tr->format_date($epoch, 'YYYY-MM-DD'); # "2026-08-09"
$tr->format_date($epoch, 'DD/MM/YYYY'); # "09/08/2026"
# Also accepts date strings
$tr->format_date("2026-08-09", 'full'); # "Pazar, 9 Ağustos 2026"
Supported tokens:
| Token | Meaning | Example |
|---|---|---|
YYYY / YY |
4/2-digit year | 2026 / 26 |
MMMM / MMM / MM / M |
Month name / short / 2-digit / single | Ağustos / Ağu / 08 / 8 |
DD / D |
Day (2-digit / single) | 09 / 9 |
dddd / ddd |
Day name / short | Pazar / Paz |
HH / H |
Hour | 14 / 14 |
mm / m |
Minute | 05 / 5 |
ss / s |
Second | 09 / 9 |
Accepted input formats:
17863452002026-08-09, 2026/08/09, 2026.08.092026-08-09 14:30:00, 2026-08-09T14:30parse_date($string [, %opts])my $epoch = $tr->parse_date("09.08.2026"); # Unix timestamp
my $h = $tr->parse_date("09.08.2026", hash => 1);
# { year => 2026, month => 8, day => 9, hour => 0, minute => 0, second => 0 }
The short date format automatically detects whether it’s
DD.MM.YYYYorMM/DD/YYYYbased on the locale’sdate_format.shortvalue.
format_number($num [, %opts])my $tr = AmberDB::Locale->new(language => "tr");
$tr->format_number(1234567.89); # "1.234.567,89"
$tr->format_number(1234567.89, decimals => 0); # "1.234.568"
$tr->format_number(1234567.89, decimals => 3); # "1.234.567,890"
my $en = AmberDB::Locale->new(language => "en");
$en->format_number(1234567.89); # "1,234,567.89"
my $fr = AmberDB::Locale->new(language => "fr");
$fr->format_number(1234567.89); # "1 234 567,89"
| Locale | Decimal | Grouping | Example |
|---|---|---|---|
tr |
, |
. |
1.234.567,89 |
en |
. |
, |
1,234,567.89 |
de |
, |
. |
1.234.567,89 |
fr |
, |
(space) | 1 234 567,89 |
ru |
, |
(space) | 1 234 567,89 |
ar |
٫ |
٬ |
١٬٢٣٤٬٥٦٧٫٨٩ |
format_currency($amount [, $code | %opts])my $tr = AmberDB::Locale->new(language => "tr");
$tr->format_currency(1234.50); # "₺1.234,50"
$tr->format_currency(1234.50, 'EUR'); # "1.234,50 €" (suffix + space)
$tr->format_currency(1234.50, currency => 'USD'); # "$1.234,50"
my $de = AmberDB::Locale->new(language => "de");
$de->format_currency(1234.50, 'EUR'); # "1.234,50 €"
Resolution priority:
%opts override → locale currencies → AmberDB::Locale::Currency universal data → default values
plural($count, \%forms)my $en = AmberDB::Locale->new(language => "en");
$en->plural(1, { one => "{count} item", other => "{count} items" });
# "1 item"
$en->plural(5, { one => "{count} item", other => "{count} items" });
# "5 items"
my $tr = AmberDB::Locale->new(language => "tr");
$tr->plural(1, { one => "{count} ürün", other => "{count} ürün" });
# "1 ürün"
$tr->plural(5, { one => "{count} ürün", other => "{count} ürün" });
# "5 ürün"
# Russian — 4 different forms
my $ru = AmberDB::Locale->new(language => "ru");
$ru->plural(1, { one => "{count} яблоко", few => "{count} яблока",
many => "{count} яблок", other => "{count} яблока" });
# "1 яблоко"
$ru->plural(3, { ... }); # "3 яблока" (few)
$ru->plural(5, { ... }); # "5 яблок" (many)
$ru->plural(11, { ... }); # "11 яблок" (many)
CLDR rule string format:
one{n==1}other
zero{n==0}one{n==1}two{n==2}few{n%100>=3&&n%100<=10}many{n%100>=11&&n%100<=99}other
one{n%10==1&&n%100!=11}few{n%10>=2&&n%10<=4&&(n%100<10||n%100>=20)}many{...}other
{count}or{n}placeholders are replaced with locale-appropriately formatted numbers (viaformat_numberwithdecimals => 0).
$lang->language(); # "tr" — active language tag
$lang->months(); # ["Ocak", "Şubat", ..., "Aralık"]
$lang->days(); # ["Pazar", "Pazartesi", ..., "Cumartesi"]
first_char($string) — Alphabetical index character$tr->first_char(" çarşı "); # "Ç"
$tr->first_char("123abc"); # "0-9"
$tr->first_char("İzmir"); # "İ"
AmberDB::Locale::Currency — Universal Currency Data12 currencies are defined in the ISO 4217 standard:
| Code | Name | Symbol | Decimals |
|---|---|---|---|
| TRY | Türk Lirası | ₺ | 2 |
| USD | ABD Doları | $ | 2 |
| EUR | Euro | € | 2 |
| GBP | İngiliz Sterlini | £ | 2 |
| RUB | Rus Rublesi | ₽ | 2 |
| AZN | Manat | ₼ | 2 |
| SAR | Suudi Riyali | ر.س | 2 |
| JPY | Japon Yeni | ¥ | 0 |
| CHF | İsviçre Frangı | CHF | 2 |
| CAD | Kanada Doları | CA$ | 2 |
| AUD | Avustralya Doları | A$ | 2 |
| CNY | Çin Yuanı | ¥ | 2 |
AmberDB::Locale::Currency->by_code('TRY'); # { num=>'949', name=>'Türk Lirası', symbol=>'₺', digits=>2 }
AmberDB::Locale::Currency->symbol('EUR'); # "€"
AmberDB::Locale::Currency->name('USD'); # "ABD Doları"
AmberDB::Locale::Currency->all(); # [ ['TRY','Türk Lirası'], ['USD','ABD Doları'], ... ]
AmberDB::Locale::Currency->active_codes(); # qw(TRY USD EUR GBP RUB AZN SAR JPY CHF CAD AUD CNY)
To add a new language, create an Amber/Locale/Lang/<code>.pm file:
package AmberDB::Locale::Lang::it; # Italian example
use strict;
use warnings;
our $VERSION = '1.0';
sub data {
return {
uc_map => {},
lc_map => {},
sort_map => {},
alphabet_chars => "\x{E0}\x{E8}\x{E9}\x{EC}\x{F2}\x{F9}", # à è é ì ò ù
accent_map => {},
ascii_map => {}, # NFD handles all accents
numbers => {
zero => 'Zero',
ones => [qw(Uno Due Tre Quattro Cinque Sei Sette Otto Nove)],
tens => [qw(Dieci Venti Trenta Quaranta Cinquanta Sessanta Settanta Ottanta Novanta)],
hundred => 'Cento',
thousand => 'Mille',
million => 'Milione',
billion => 'Miliardo',
currency => { main => 'EUR', sub => 'centesimo' },
decimal_sep => ',',
hundred_one_prefix => 0,
thousand_one_prefix => 0,
},
html_entities => {},
months => [qw(Gennaio Febbraio Marzo Aprile Maggio Giugno
Luglio Agosto Settembre Ottobre Novembre Dicembre)],
days => [qw(Domenica Lunedì Martedì Mercoledì Giovedì Venerdì Sabato)],
number_format => {
decimal_sep => ',',
group_sep => '.',
group_size => 3,
},
default_currency => 'EUR',
currency_position => 'suffix',
currency_space => 1,
date_format => {
short => 'DD/MM/YYYY',
medium => 'D MMM YYYY',
long => 'D MMMM YYYY',
full => 'dddd D MMMM YYYY',
time => 'HH:mm',
datetime => 'DD/MM/YYYY HH:mm',
},
plural_rule => 'one{n==1}other',
};
}
1;
The file name is automatically
required by_load_locale()inAmberDB::Locale. No additional registration is needed.
| Topic | Detail |
|---|---|
| Instance cache | Only 1 instance is created per language; subsequent new() calls return from cache |
| Regex pre-compile | All patterns (_uc_re, _lc_re, _sort_re, _accent_re, _ascii_re, _safe_re) are compiled at construction time; not recompiled on each method call |
| Sort key length | uc_map/lc_map keys are sorted by length descending → multi-character mappings match first |
| Unicode::Collate | May be slow on first load (table reading), but subsequent calls are fast |
| Problem | Cause | Solution |
|---|---|---|
i → I instead of İ |
Using en locale |
Provide language => "tr" |
to_ascii output has a instead of ae |
In en locale, NFD does ä → a |
Use de locale (DIN 5007-2: ä → ae) |
| Number text returns empty | Input contains only separators/punctuation | Check for valid digit input |
| Eastern Arabic digits not converted | normalize_num not called explicitly |
num2text/format_number does it automatically; no manual call needed |
| Unknown language error | Lang module file doesn’t exist | Create Amber/Locale/Lang/<code>.pm or accept en fallback |
my $L = AmberDB::Locale->new(language => "tr");
# Text transformations
$L->uc("ığdır") # IĞDIR
$L->lc("İSTANBUL") # istanbul
$L->ucfirst("merhaba dünya") # Merhaba Dünya
$L->fold("İSTANBUL") # istanbul
$L->ieq("İstanbul", "istanbul") # 1
$L->normalize("<b>Kâr</b> &") # Kar &
$L->to_ascii("çarşı") # carsi
$L->to_ascii("çarşı", 1) # carsi (slug)
$L->first_char("çarşı") # Ç
$L->substring("Şanlıurfa", 0, 5) # Şanlı
# Sorting
$L->sort(["İzmir","Ankara","Van"]) # Ankara, İzmir, Van
# Numbers
$L->num2text(1234.56) # Bin İki Yüz Otuz Dört TL Elli Altı KR
$L->format_number(1234567.89) # 1.234.567,89
$L->format_currency(99.9, 'TRY') # ₺99,90
# Date
$L->format_date(time, 'full') # Pazar, 9 Ağustos 2026
$L->parse_date("09.08.2026") # epoch
# Plural
$L->plural(1, {one=>"{count} adet", other=>"{count} adet"}) # 1 adet
# Accessors
$L->language() # tr
$L->months() # [Ocak, Şubat, ...]
$L->days() # [Pazar, Pazartesi, ...]