Bulgarian Pronunciation and IPA Transcription Guide
← Back to Bulgarian transcription ↑ All Languages
- No stress mark, word with two or more vowels: every vowel is treated as unstressed and reduced (а, ъ → ɐ; о, у → o), and no ˈ is printed: добър /dobɐr/. Only ɛ and i are left alone.
- With the acute accent: the marked vowel keeps its full quality and the stress mark is placed at the start of the syllable: до́бър /ˈdɔbɐr/.
- Word with a single vowel and no mark is not reduced: кон /kɔn/, хляб /xlʲa̟p/.
Table of Contents
- Getting Started
- Quick Reference (Glossary)
- How to Read IPA Symbols
- Bulgarian Pronunciation Guide
- Orthography to IPA Mapping
- Using the Tool: Practical Notes
- For Developers
- Related Pronunciation Guides
About This Tool
This Bulgarian transcription app uses the Wiktionary Bulgarian Pronunciation Module (the local copy is
lua_modules/bg-pron_wasm.lua) to generate phonemic transcriptions for Bulgarian text. IPA
(International Phonetic Alphabet) is a standardized system for transcribing speech sounds using symbols based on
the Latin alphabet. This tool converts Bulgarian Cyrillic spelling into IPA, helping learners, linguists, and
developers understand pronunciation.
The conversion is purely rule-based: the page calls window.bg_ipa.toIPA(cleanText)
and nothing else. There is no lexicon and no stress dictionary, so the result is only as good
as the input you give it. The rules cover letter-to-sound mapping, the movement of the stress mark to a syllable
onset, vowel reduction around the stress, palatalization before я/ю/ь, voicing assimilation, final devoicing,
hard ɫ, nasal assimilation and a few consonant-cluster simplifications. They are
described step by step in Implementation Details.
Dialects & Limitations
- Supported: Standard (literary) Bulgarian, one style: Default.
- Not supported: Regional dialects; the output describes the standard literary pronunciation only.
- No automatic stress: stress must be typed with a combining acute accent; see the box at the top of this page.
- Edge cases: Loanwords with non-standard sounds, words where a stressed а/я in an ending is really pronounced ɤ, and proper names are not handled specially.
- Input cleaning: before transcription the app removes everything except letters, combining marks, hyphens and apostrophes (see Spaces, Hyphens and Unsupported Characters).
Phonemic vs Phonetic
- Phonemic (the only form offered): a broad transcription that shows the sounds that distinguish words. In practice the module is a little more detailed than a pure phonemic transcription: it marks the reduced vowel ɐ, the fronted vowels after soft consonants (a̟), hard ɫ and the nasal allophones ŋ, ɱ.
- Phonetic: not available for Bulgarian. The handler returns nothing for any other form
(
if (langForm !== "Phonemic") return null;).
Quick Reference (Glossary)
- Phoneme
- The smallest unit of sound in a language that distinguishes meaning.
- Allophone
- A variant pronunciation of a phoneme that doesn't change meaning, e.g. the hard ɫ versus the plain l.
- Stress
- Bulgarian stress is free (can fall on any syllable) and is not marked in ordinary spelling. This tool needs you to supply it with a combining acute accent after the vowel. Shown in IPA as ˈ placed before the stressed syllable: ма́ма /ˈmamɐ/.
- Vowel Reduction
- Unstressed vowels become weaker and more central. Here: а and ъ become ɐ, о and у become o: вода́ /voˈda/.
- Schwa-like Vowel (ъ)
- The letter ъ is the Bulgarian mid-central vowel ɤ when stressed (зъ́б /ˈzɤp/) and ɐ when unstressed.
- Palatalization
- Softening of a consonant before я, ю or ь, written ʲ: ля́ва /ˈlʲa̟vɐ/, ми́ля /ˈmilʲɐ/.
- Fronted Vowel
- A back vowel moved forward after a soft or postalveolar consonant, written with the advanced diacritic ◌̟: ша́пка /ˈʃa̟pkɐ/.
- Affricate
- A stop released as a fricative, written with a tie bar: t͡s (ц), t͡ʃ (ч): ци́рк /ˈt͡sirk/, чо́век /ˈt͡ʃɔvɛk/.
- Tie Bar
- The combining mark ͡ joining the two halves of an affricate into one sound.
- Voicing Assimilation
- A consonant takes the voicing of the consonant after it: про́сба /ˈprɔzbɐ/ (с before б becomes z), ска́зка /ˈskaskɐ/ (з before к becomes s).
- Final Devoicing
- Voiced obstruents are pronounced voiceless at the end of a word: гра́д /ˈɡrat/, хле́б /ˈxlɛp/.
- Velar Nasal
- ŋ, the "ng" of English "sing", used for н before к or г: ба́нка /ˈbaŋkɐ/.
- Labiodental Nasal
- ɱ, used for м before ф or в: ка́мфор /ˈkaɱfor/.
- Hard L
- ɫ, the "dark l" used before back vowels and consonants and at the end of a word: хоте́л /xoˈtɛɫ/. Before е, и or a soft sign the plain l is kept: ле́ка /ˈlɛkɐ/.
- Combining Acute (U+0301)
- The character you type after a vowel to mark primary stress. Mapped internally to ˈ.
- Combining Grave (U+0300)
- Marks secondary stress (mapped to ˌ). Only allowed in a word that also has an acute accent.
- Dot Below (U+0323)
- Typed under а or я to force the sound ɤ (as if it were ъ): пътя̣́т /pɐˈtʲɤ̟t/, compare пътя́т /pɐˈtʲa̟t/.
- Clitic
- A short word that leans on its neighbour. Single-letter prepositions в and с are transcribed alone as в /f/ and с /s/.
- NFD
- Unicode Normalization Form Decomposed. The module decomposes the input so that accents become separate characters it can see.
- Obstruent / Sonorant
- Obstruents (stops, fricatives, affricates) can be voiced or voiceless. Sonorants (m n l r j) are always voiced and never trigger voicing assimilation.
- Syllable Onset
- The consonants at the start of a syllable. The stress mark is placed before the onset, not directly before the vowel: the mark goes in front of the whole onset (see гра́д /ˈɡrat/), not directly before the vowel.
How to Read IPA Symbols
This section lists the symbols that the engine actually emits for Bulgarian (collected by running a broad word list through it). For each symbol there is an example and an approximate English equivalent.
Vowel Symbols
| IPA | Example | English Approximation | Notes |
|---|---|---|---|
| a | ба́ба /ˈbabɐ/ | "father" | Stressed а, also stressed я after a soft consonant |
| ɐ | вода́ /voˈda/ | "cut" / "sofa" | Unstressed а and unstressed ъ (the reduced vowel) |
| ɛ | ве́чер /ˈvɛt͡ʃɛr/ | "bed" | е in every position. Never reduced by this module |
| i | ми́р /ˈmir/ | "machine" | и in every position. Never reduced by this module |
| ɔ | мо́ре /ˈmɔrɛ/ | "thought" (UK) | Stressed о |
| o | рабо́та /rɐˈbɔtɐ/ | "go" (without glide) | Unstressed о and unstressed у |
| u | ту́ка /ˈtukɐ/ | "food" | Stressed у (and ю after a consonant, with ʲ) |
| ɤ | зъ́б /ˈzɤp/ | No English equivalent; "bird" without r-colouring | Stressed ъ |
Consonant Symbols
| IPA | Example | English Approximation | Notes |
|---|---|---|---|
| p b t d k ɡ | па́рче /ˈpart͡ʃɛ/, ба́ба /ˈbabɐ/ | "spin", "bad", "stop", "do", "sky", "go" | Plain stops (unaspirated). The symbol ɡ is the IPA single-storey g. |
| f v s z | ва́жен /ˈvaʒɛn/, со́фия /ˈsɔfijɐ/ | "five", "van", "see", "zoo" | Subject to voicing assimilation and final devoicing |
| ʃ | ша́пка /ˈʃa̟pkɐ/ | "shoe" | ш; also the first half of щ |
| ʒ | жа́ба /ˈʒa̟bɐ/ | "measure" | ж |
| x | хле́б /ˈxlɛp/ | Scottish "loch" | х. Voiced to ɣ before б, д, г, з, ж |
| ɣ | бе́хдом /ˈbɛɣdom/ | No English equivalent | Voiced allophone of х, only before a voiced obstruent |
| t͡s | ци́рк /ˈt͡sirk/ | "cats" | ц, affricate |
| t͡ʃ | чо́век /ˈt͡ʃɔvɛk/ | "church" | ч, affricate |
| ʃt | ща́стие /ˈʃtastiɛ/ | "fish tank" | щ is written as two sounds, with no tie bar |
| m n | ма́ма /ˈmamɐ/, не́бе /ˈnɛbɛ/ | "man", "no" | Plain nasals |
| ŋ | ба́нка /ˈbaŋkɐ/ | "sing" | н before к or г |
| ɱ | ко́мфорт /ˈkɔɱfort/ | "symphony" (the m before f) | м before ф or в |
| l | ле́ка /ˈlɛkɐ/ | "leaf" | Before е, и and ʲ |
| ɫ | ма́ла /ˈmaɫɐ/ | "full" (dark l) | л elsewhere |
| r | ра́бота /ˈrabotɐ/ | Spanish "pero" (tap or trill) | р |
| j | ма́йка /ˈmajkɐ/, ю́ли /ˈju̟li/ | "yes" | й; also я/ю/ь after a vowel or at the start of a word |
| ʲ | ми́ля /ˈmilʲɐ/ | The "y" glide in "million" | Softness (palatalization) of the preceding consonant |
| w | ўт /wt/ | "water" | Only for the non-standard letter ў; not a regular Bulgarian letter |
Diacritical Marks
| Symbol | Name | Meaning | Example |
|---|---|---|---|
| ˈ | Primary stress | Main emphasis; placed before the stressed syllable's onset. Produced from the combining acute U+0301 | ра́дио /ˈradio/ |
| ˌ | Secondary stress | Weaker emphasis. Produced from the combining grave U+0300, only valid together with an acute | ма̀ма́ /ˌmaˈma/ |
| ʲ | Palatalization | The consonant before it is soft | ля́ва /ˈlʲa̟vɐ/ |
| ◌̟ | Advanced | Vowel moved forward after ʃ, ʒ, ʲ or j | ша́рлатан /ˈʃa̟rɫɐtɐn/ |
| ◌͡◌ | Tie bar | Two segments form one affricate | ча́й /ˈt͡ʃa̟j/ |
Interactive Features
- Click words to cycle variants: the Bulgarian module returns exactly one transcription per word, so there is nothing to cycle through.
- Audio playback: Click the speaker icon next to any word or line to hear text-to-speech pronunciation (requires browser TTS support). The browser or Edge voice reads the Cyrillic text, not the IPA.
- Export results: Use PDF or CSV buttons to save transcriptions in your preferred format.
Multiple Pronunciation Variants
Bulgarian produces one transcription per input. If a word can be stressed in more than one way (for example homographs such as за́мък "castle" and замъ́к "lock"), you get a different result by typing the stress mark in a different position; the tool cannot offer both for you to cycle through.
Bulgarian Pronunciation Guide
Bulgarian spelling is close to phonemic: almost every letter has one sound. The exceptions are the free stress (which is not written), the reduction of unstressed vowels, and a set of voicing and softening rules. This guide explains what the engine does, in the order a reader meets the rules.
Typing the Stress
Because Bulgarian dictionaries mark stress and ordinary text does not, the module reads the stress from the input:
- Acute accent U+0301 after the vowel: ма́ма /ˈmamɐ/ and мама́ /mɐˈma/ have the same letters but different stress and therefore different vowels: the first vowel of the second word is reduced.
- Where the mark ends up: the accent is moved to the beginning of the syllable. It jumps over the vowel, over a soft sign or й, and over a single consonant, plus the clusters described in stage 7: ра́бота /ˈrabotɐ/, сла́дко /ˈsɫatko/, ска́зка /ˈskaskɐ/, хле́б /ˈxlɛp/.
- Word-initial consonants: if only consonants precede the stressed vowel, the mark goes to the very start of the word: гра́д /ˈɡrat/, хле́б /ˈxlɛp/.
- Prefixes: the prefixes без-, въз-, из-, раз- and a few others keep the stress mark after the prefix: безбра́чие /bɛzˈbrat͡ʃiɛ/, разка́з /rɐsˈkas/.
- Two stress marks: allowed and both are printed: ма́ма́ /ˈmaˈma/. Secondary stress with the grave accent: ма̀ма́ /ˌmaˈma/.
Vowels and Unstressed Reduction
The module reduces vowels around the stress, not just before it:
- Before the stress: everything before the stress mark is reduced: a, ɤ → ɐ and ɔ, u → o: вода́ /voˈda/, ръка́ /rɐˈka/.
- After the stress: every vowel after the stressed one is reduced in the same way: ма́ма /ˈmamɐ/, ра́бота /ˈrabotɐ/, ма́ймуна /ˈmajmonɐ/.
- Never reduced: ɛ and i stay as they are in all positions: кафе́ /kɐˈfɛ/, ми́лион /ˈmilion/.
- No stress mark and more than one vowel: the whole word counts as "before the stress", so all vowels are reduced: добър /dobɐr/ against до́бър /ˈdɔbɐr/. A word with exactly one vowel letter is never reduced: кон /kɔn/, съм /sɤm/.
- The vowel count is made on the whole input (letters а ъ о у е и ѝ ю я), not per word. If you paste two words into one call, two vowels are enough to reduce both.
Soft Consonants: я, ю, ь
- я = ʲa and ю = ʲu. After a consonant the ʲ softens that consonant: ля́ва /ˈlʲa̟vɐ/, лю́бов /ˈlʲu̟bof/, ми́ля /ˈmilʲɐ/.
- After a vowel or at the beginning of a word the glide becomes a full j: ю́ли /ˈju̟li/, я́вор /ˈja̟vor/, ма́я /ˈmajɐ/, ра́дио /ˈradio/.
- ь is written only before о (ьо) in Bulgarian. After a consonant it gives ʲ: ка́ньон /ˈkanʲo̟n/, шофьо́р /ʃo̟ˈfʲɔr/; after a vowel it gives j: ма́ьо /ˈmajo̟/.
- Fronting: the vowels a, o, u, ɤ immediately after ʃ, ʒ, ʲ, j get the advanced diacritic: ша́пка /ˈʃa̟pkɐ/, жа́ба /ˈʒa̟bɐ/, ча́й /ˈt͡ʃa̟j/ (the ʃ of ч counts), ля́ва /ˈlʲa̟vɐ/. The stressed ɔ and the reduced ɐ are not fronted.
- Unstressed я/ю: they reduce along with the vowel: ми́ля /ˈmilʲɐ/ ends in ʲɐ; любо́в /lʲo̟ˈbɔf/ gives ʲo̟.
Ц, Ч, Щ, ДЖ and ДЗ
- ц = t͡s, ч = t͡ʃ (with tie bar): ци́рк /ˈt͡sirk/, чо́век /ˈt͡ʃɔvɛk/.
- щ = ʃt (two sounds, no tie bar): ща́стие /ˈʃtastiɛ/, пе́щ /ˈpɛʃt/. Because щ is expanded before the later rules run, the t can be lost in clusters (see Consonant Cluster Simplification): по́щна /ˈpɔʃnɐ/.
- дж and дз are not treated as single affricate letters by the transcription rules. They are simply д + ж and д + з: дже́м /ˈdʒɛm/, джо́б /ˈdʒɔp/, дзи́ндзър /ˈdzindzɐr/. The result has no tie bar, unlike ц and ч. Voicing assimilation and final devoicing apply to д and ж separately: хо́дж /ˈxɔtʃ/.
Voicing Assimilation
A group of voiced obstruents (б в г д ж з) before a voiceless one (п т к с ш ф х) or at the end of the word becomes voiceless; a voiceless obstruent before б, д, г, з, ж becomes voiced.
| Process | Pairs | Examples |
|---|---|---|
| Regressive devoicing (voiced before voiceless) | b→p d→t ɡ→k z→s ʒ→ʃ v→f | ска́зка /ˈskaskɐ/, ре́дко /ˈrɛtko/, вто́рник /ˈftɔrnik/, сле́дващ /ˈslɛdvɐʃt/ |
| Regressive voicing (voiceless before б д г з ж) | p→b t→d k→ɡ s→z ʃ→ʒ x→ɣ f→v | про́сба /ˈprɔzbɐ/, се́здесет /ˈsɛzdɛsɛt/, бе́хдом /ˈbɛɣdom/ |
| Voiced consonant before a sonorant or vowel | unchanged | тра́вма /ˈtravmɐ/, ра́жда /ˈraʒdɐ/ |
The letter в is devoiced like б, г, д, ж, з: вто́рник /ˈftɔrnik/ (в before т). Unlike them it does not trigger voicing in the consonant before it: ква́рта /ˈkvartɐ/ keeps its к.
Final Devoicing
At the end of a word, voiced obstruents are devoiced. The same rule that does voicing assimilation handles it,
because the word-boundary symbol # is in the "voiceless" set.
| Spelling | Final | Elsewhere | Example |
|---|---|---|---|
| б | p | b | хле́б /ˈxlɛp/ |
| в | f | v | ръка́в /rɐˈkaf/, лю́бов /ˈlʲu̟bof/ |
| г | k | ɡ | кни́га /ˈkniɡɐ/ (medial), ю́г /ˈju̟k/ |
| д | t | d | гра́д /ˈɡrat/ |
| ж | ʃ | ʒ | гара́ж /ɡɐˈraʃ/, гара́жът /ɡɐˈraʒɐt/ |
| з | s | z | зъ́б /ˈzɤp/ (initial з), ска́з /ˈskas/ |
Hard L
The letter л is ɫ unless the next symbol is ɛ, i or the softness mark ʲ. This includes the end of a word and l before another consonant:
- Plain l: ле́ка /ˈlɛkɐ/, ли́ст /ˈlist/, ми́ля /ˈmilʲɐ/, лю́лка /ˈlʲu̟ɫkɐ/.
- Dark ɫ: ма́ла /ˈmaɫɐ/, хоте́л /xoˈtɛɫ/, ма́лко /ˈmaɫko/, сла́дко /ˈsɫatko/, ко́лело /ˈkɔlɛɫo/.
Nasal Assimilation: ŋ and ɱ
- н → ŋ before к or г: ба́нка /ˈbaŋkɐ/, тъ́нко /ˈtɤŋko/, ме́нгеме /ˈmɛŋɡɛmɛ/.
- м → ɱ before ф or в: ка́мфор /ˈkaɱfor/, ко́мвинат /ˈkɔɱvinɐt/.
Consonant Cluster Simplification
- Sibilant assimilation: с/з before ш, ж or the ш of щ and ч becomes the same sibilant as the one that follows: безшу́мен /bɛʃˈʃu̟mɛn/, разшири́ /rɐʃʃiˈri/, изчи́стя /iʃˈt͡ʃistʲɐ/, изжи́вея /iʒˈʒivɛjɐ/.
- Loss of т or д after a sibilant: a т/д after с, з, ш, ж and before т, д, к, н, м, л is dropped: че́стно /ˈt͡ʃɛsno/, по́здно /ˈpɔzno/, ра́здна /ˈraznɐ/, ща́стлив /ˈʃtaslif/. Since щ = ʃt, the same applies to щ: по́щна /ˈpɔʃnɐ/ but по́щенски /ˈpɔʃtɛnski/ (the vowel е blocks it).
- Double consonants are kept as written: кла́сс /ˈkɫass/.
Spaces, Hyphens and Unsupported Characters
- Spaces and punctuation are deleted by the app before the module sees the text (only letters, combining marks, hyphen, apostrophe survive). Two words in one call become one word: до́бър де́н /ˈdɔbɐrˈdɛn/. Transcribe one word at a time for meaningful results.
- Hyphens are kept in the output and are not treated as word boundaries: вя́тър-ма́йстор /ˈvʲa̟tɐr-ˈmajstor/. Stress is processed for the whole string, so a hyphenated compound with no mark at all is reduced in both halves: добре-дошли /dobrɛ-doʃli/.
- Capitals are lowercased, so Бурга́с /borˈɡas/ and бурга́с give the same result.
- Letters that are not in the mapping pass through unchanged to the output: Latin letters (hello /heɫɫo/), Russian ы (ды́лбя), digits are removed by the cleaning step.
Comprehensive Spelling-to-IPA Mapping
These tables reflect the rules the engine applies. Letters are listed with their basic mapping and the changes caused by context.
1. Vowels
| Letter | Stressed | Example (stressed) | Unstressed | Example (unstressed) |
|---|---|---|---|---|
| а | a | ба́ба /ˈbabɐ/ | ɐ | ма́ма /ˈmamɐ/ |
| ъ | ɤ | гъ́ска /ˈɡɤskɐ/ | ɐ | къща́та /kɐʃˈtatɐ/ |
| о | ɔ | по́ле /ˈpɔlɛ/ | o | вода́ /voˈda/ |
| у | u | су́мка /ˈsumkɐ/ | o | уча́т /oˈt͡ʃa̟t/ |
| е | ɛ | ме́со /ˈmɛso/ | ɛ | нейно́ /nɛjˈnɔ/ |
| и | i | ни́ /ˈni/ | i | интерне́т /intɛrˈnɛt/ |
| я | ʲa̟ | тя́ /ˈtʲa̟/ | ʲɐ | ми́ля /ˈmilʲɐ/ |
| ю | ʲu̟ | лю́лка /ˈlʲu̟ɫkɐ/ | ʲo̟ | любо́в /lʲo̟ˈbɔf/ |
Word-initially and after a vowel, я and ю start with j instead of ʲ: я́вор /ˈja̟vor/, ю́жен /ˈju̟ʒɛn/, ма́я /ˈmajɐ/.
2. Single Consonant Letters
| Letter | Basic IPA | Example | Notes |
|---|---|---|---|
| б | b | ба́ба /ˈbabɐ/ | Final: p |
| в | v | ва́рна /ˈvarnɐ/ | Final or before voiceless: f |
| г | ɡ | го́ра /ˈɡɔrɐ/ | Final: k |
| д | d | да́ма /ˈdamɐ/ | Final: t |
| ж | ʒ | жи́вот /ˈʒivot/ | Final: ʃ |
| з | z | зъ́б /ˈzɤp/ | Final: s |
| й | j | йо́гурт /ˈjɔɡort/ | Also in the diphthong-like sequences: ни́кой /ˈnikoj/ |
| к | k | ку́пя /ˈkupʲɐ/ | Before б д г з ж: ɡ |
| л | ɫ / l | ма́лко /ˈmaɫko/, ле́ка /ˈlɛkɐ/ | l only before ɛ, i, ʲ |
| м | m | ма́ма /ˈmamɐ/ | Before ф/в: ɱ |
| н | n | не́бе /ˈnɛbɛ/ | Before к/г: ŋ |
| п | p | па́рче /ˈpart͡ʃɛ/ | Before б д г з ж: b |
| р | r | ра́бота /ˈrabotɐ/ | No variants |
| с | s | се́дем /ˈsɛdɛm/ | Before б д г з ж: z |
| т | t | то́й /ˈtɔj/ | Before б д г з ж: d |
| ф | f | ко́фа /ˈkɔfɐ/ | Before б д г з ж: v |
| х | x | хле́б /ˈxlɛp/ | Before б д г з ж: ɣ |
| ш | ʃ | ша́пка /ˈʃa̟pkɐ/ | Before б д г з ж: ʒ |
3. Digraphs and Composite Letters
| Spelling | IPA | Example | Notes |
|---|---|---|---|
| ц | t͡s | ци́гани /ˈt͡siɡɐni/ | Affricate with tie bar |
| ч | t͡ʃ | чи́чо /ˈt͡ʃit͡ʃo̟/ | Affricate with tie bar |
| щ | ʃt | ща́м /ˈʃtam/ | Two sounds; the t may drop in clusters |
| дж | dʒ | джа́ма /ˈdʒa̟mɐ/ | д + ж, no tie bar |
| дз | dz | дзвъ́нец /ˈdzvɤnɛt͡s/ | д + з, no tie bar |
| ьо | ʲɔ / jɔ | ньо́ки /ˈnʲɔki/, ко́ньо /ˈkɔnʲo̟/ | ь is ʲ after a consonant, j after a vowel or word-initially |
4. Context-Dependent Rules
| Context | Result | Example |
|---|---|---|
| Voiced obstruent before voiceless obstruent | devoiced | ска́зка /ˈskaskɐ/ |
| Voiceless obstruent before б д г з ж | voiced | про́сба /ˈprɔzbɐ/ |
| Word-final voiced obstruent | devoiced | гра́д /ˈɡrat/ |
| н before к, г | ŋ | ба́нка /ˈbaŋkɐ/ |
| м before ф, в | ɱ | ка́мфор /ˈkaɱfor/ |
| с/з before ш, ж, ч | copies the following sibilant | разшири́ /rɐʃʃiˈri/ |
| с/з/ш/ж + т/д + т д к н м л | the т/д is dropped | че́стно /ˈt͡ʃɛsno/ |
| а, о, у, ъ after ш, ж, ʲ, j | advanced (◌̟) | жа́ба /ˈʒa̟bɐ/ |
5. Stress Marks and Where They Land
| Input pattern | Mark placed | Example |
|---|---|---|
| V́ preceded by one consonant | before that consonant | ма́ма /ˈmamɐ/ |
| V́ preceded by obstruent + л/р | before the pair (e.g. хл, кр) | хле́б /ˈxlɛp/, кра́ва /ˈkravɐ/ |
| V́ preceded by к/г + в | before the pair | ква́рта /ˈkvartɐ/ |
| V́ preceded by с/з + stop or sonorant | before the pair | сла́дко /ˈsɫatko/, ска́зка /ˈskaskɐ/ |
| V́ preceded only by consonants | word start | гра́д /ˈɡrat/ |
| V́ after a prefix such as без-, из-, раз- | after the prefix | безбра́чие /bɛzˈbrat͡ʃiɛ/ |
| Grave + acute | ˌ and ˈ | ма̀ма́ /ˌmaˈma/ |
6. Palatalization and Fronting
| Input | Rule | Example |
|---|---|---|
| consonant + я | ʲa (a advanced if stressed) | дя́до /ˈdʲa̟do/ |
| consonant + ю | ʲu (u advanced if stressed) | лю́бен /ˈlʲu̟bɛn/ |
| word start / vowel + я, ю | j + vowel | я́вор /ˈja̟vor/, ма́я /ˈmajɐ/ |
| ʃ, ʒ + a, o, u, ɤ | advanced vowel | шо́фьор /ˈʃɔfʲo̟r/ |
| consonant + ь + о | ʲ + о | ка́ньон /ˈkanʲo̟n/ |
Using the Tool: Practical Notes
How do I mark the stress, and what happens if I do not?
Type a combining acute accent (U+0301) directly after the stressed vowel: до́бър /ˈdɔbɐr/. In a word with more than one vowel and no mark, every vowel is reduced and no stress sign is printed; a word with a single vowel is not reduced: кон /kɔn/. See Typing the Stress.
Why is there only one transcription for each word?
The module returns one result per input and consults no dictionary. When homographs differ in stress, such as за́мък /ˈzamɐk/ and замъ́к /zɐˈmɤk/, the position of the mark you type selects the reading. See Multiple Pronunciation Variants.
Why can the output differ from a dictionary transcription?
The rules work from the spelling and the typed stress alone. For example, the stressed ending of third-person plural verbs is pronounced with ɤ, which the engine produces only if a dot below (U+0323) is typed under the vowel: говоря́т /ɡovoˈrʲa̟t/, говоря̣́т /ɡovoˈrʲɤ̟t/. Dialects, loanword exceptions and proper names are not handled specially. See Common Issues & Limitations. Open in the converter
Is the transcription phonemic or phonetic?
Only a phonemic form is offered. It is a broad transcription that nevertheless marks a few predictable details: the reduced vowel ɐ, vowels fronted after soft consonants, dark ɫ and the nasals ŋ and ɱ. See Phonemic vs Phonetic.
Implementation Details (for Developers)
Everything the app uses lives in the single function export.toIPA(term, endschwa) of
lua_modules/bg-pron_wasm.lua. It is a chain of Lua string substitutions (rsub is a
wrapper over mw.ustring.gsub that returns only the string; rsub_repeatedly repeats
until nothing changes). The tables phonetic_chars_map, devoicing,
voicing and IPA_prefixes drive the substitutions. The stages below follow the code
in execution order.
- Input cleaning (JavaScript, before Lua):
sanitize()inscripts/utils.jsremoves everything except\p{L},\p{M}, apostrophes and hyphens, then NFKC-normalizes. The Bulgarian handler callswindow.bg_ipa.toIPA(cleanText)only when the form is Phonemic. - Normalization: the term is lowercased and decomposed (NFD); у + breve is
recomposed to
ўand и + breve toй. - Grave-accent check: a grave without any acute raises a Lua error.
- Dot-below shortcut:
а + U+0323becomesъ, andя + U+0323becomesʲɤ; this is the same as theendschwaoption, but written in the input. - Letter mapping: every character is replaced via
phonetic_chars_map. - Word boundaries: whitespace runs are wrapped in
#and the term itself gets a#at both ends. - endschwa: if the Lua argument is set, a stressed final
a(ˈ)t?#becomesɤ. - Soft sign normalization:
ʲafter a vowel or#becomesj. - Stress movement: the accents are moved leftward in a fixed series of substitutions (see the table).
- Vowel reduction: two substitutions with
reduce_vowel(). - Fronting:
([ʃʒʲj])([aouɤ])gets U+031F. - Hard L:
rsub_repeatedly("l([^ʲɛi])", "ɫ%1"). - Voicing assimilation and nasals: devoicing, voicing, ŋ, ɱ.
- Sibilant assimilation.
- Cluster reduction.
- Strip
#: the result is returned.
Stage 1-2: Input Cleaning and Normalization
| Rule | Example Input | Output | Description |
|---|---|---|---|
sanitize() |
до́бър де́н | до́бър де́н /ˈdɔbɐrˈdɛn/ | The space is removed before Lua is called, so the two words are processed as one. |
mw.ustring.lower |
Бурга́с | Бурга́с /borˈɡas/ | Capital letters become lowercase. |
| NFD + recompose й, ў | ча́й | ча́й /ˈt͡ʃa̟j/ | The breve is rejoined with и or у so that й and ў are single symbols. |
| Grave check | ма̀ма | ма̀ма /ма̀ма/ | error("Use acute accent, not grave accent, for primary stress") is raised; the handler
then returns the input unchanged (observed). The precomposed ѝ decomposes to и + grave and fails the
same way.
|
| Dot below | пътя̣́т | пътя̣́т /pɐˈtʲɤ̟t/ | я + U+0323 becomes ʲɤ; compare пътя́т /pɐˈtʲa̟t/ with a plain я. |
Stage 5: Character Mapping (phonetic_chars_map)
| Rule | Example Input | Output | Description |
|---|---|---|---|
| One-to-one letters | а б в г д е ж з и й к л м н о п р с т у ф х ш ъ | a b v ɡ d ɛ ʒ z i j k l m n ɔ p r s t u f x ʃ ɤ | Straight table lookup. ў becomes w. |
| Affricates and щ | ц ч щ | t͡s t͡ʃ ʃt | ц and ч include the tie bar U+0361; щ is two sounds. |
| Soft vowels | я ю ь | ʲa ʲu ʲ | Softness is a separate segment ʲ placed before the vowel. |
| Accents | U+0301, U+0300 | ˈ ˌ | The mark is first emitted after the vowel (where it was typed), then moved in stage 9. |
| Not mapped | Latin letters, ы | unchanged | Characters not in the table fall through as they are: hello /heɫɫo/. |
Stage 8: Soft Sign Normalization
| Rule | Example Input | Output | Description |
|---|---|---|---|
([vowels#]accent?)ʲ → %1j |
ю́ли | ю́ли /ˈju̟li/ | ʲ at the start of a word or after a vowel becomes a real j. |
| same | ма́я | ма́я /ˈmajɐ/ | After a vowel the я gives ja. |
| not applied after a consonant | ля́ва | ля́ва /ˈlʲa̟vɐ/ | The ʲ stays and softens the л. |
Stage 9: Stress Movement
The accent is moved leftward by a series of substitutions, each applied once to the whole string. The numbers below follow the order in the source (the tie-bar repair is folded into 7):
| Rule | Example Input | Output | Description |
|---|---|---|---|
| 1. over the vowel | ма́ма | ма́ма /ˈmamɐ/ | (vowel)(accent) → accent vowel |
2. over j or ʲ |
ля́ва | ля́ва /ˈlʲa̟vɐ/ | The mark moves in front of the soft sign. |
| 3. over a single consonant | ра́бота | ра́бота /ˈrabotɐ/ | The mark lands before the consonant (CV onset). |
4. over obstruent + l/r |
хле́б | хле́б /ˈxlɛp/ | Cl/Cr clusters with C in bdɡptkxfv: the whole cluster is the onset. |
5. over kv/ɡv |
ква́рта | ква́рта /ˈkvartɐ/ | The cluster is treated as one onset. |
| 6. over s/z + C | сла́дко | сла́дко /ˈsɫatko/ | C in bdɡptkvlrmn. Example with a stop: ска́зка /ˈskaskɐ/. |
7. over t͡s, t͡ʃ (ц, ч) or d + sibilant before a vowel |
чо́век | чо́век /ˈt͡ʃɔvɛk/ | [td]͡? + [szʃʒ] + vowel or ʲ: the affricate counts as one onset. If the mark still ends up
inside a tied affricate it is moved to its right. Another example: ци́рк /ˈt͡sirk/. Because дж has no tie bar the
same pattern keeps it together: джо́б /ˈdʒɔp/.
|
| 8. word-initial consonants | гра́д | гра́д /ˈɡrat/ | #(C*)(accent) → #accent C*: the mark goes to the start of the word. |
9. IPA_prefixes correction |
безбра́чие | безбра́чие /bɛzˈbrat͡ʃiɛ/ | Prefixes bɛz, vɤz, vɤzproiz, iz, naiz, poiz, prɛvɤz, proiz, raz: if the mark moved into the
prefix's last consonant, it is moved back after it. Another example: разка́з /rɐsˈkas/.
|
Explicit . boundary |
(not reachable) | - | Code supports a dot typed in the consonant cluster to force the syllable break, then removes any dots. The app's cleaning step strips dots, so only direct calls to the Lua function can use it. |
Stage 10: Vowel Reduction
| Rule | Example Input | Output | Description |
|---|---|---|---|
reduce_vowel map |
a ɔ ɤ u | ɐ o ɐ o | The table {a="ɐ", ɔ="o", ɤ="ɐ", u="o"}. ɛ and i are not in it. |
Before the accent: (#[^#accents]*)(.-#) |
вода́ | вода́ /voˈda/ | Everything from the word start up to the first accent is reduced. If the word has no accent at all this
is the whole word, unless count_vowels(origterm) <= 1: добър /dobɐr/ against the
single-vowel кон /kɔn/. The FIXME in the source points out that the single-vowel test was meant for
monosyllables.
|
After the accent: (accent[^aɛiɔuɤ#]*[aɛiɔuɤ])([^#accents]*) |
ра́бота | ра́бота /ˈrabotɐ/ | The first vowel after the accent (the stressed one) is skipped; every following vowel up to the next accent is reduced. |
Stage 11-12: Fronting and Hard L
| Rule | Example Input | Output | Description |
|---|---|---|---|
([ʃʒʲj])([aouɤ]) → %1%2U+031F |
ша́пка | ша́пка /ˈʃa̟pkɐ/ | Runs after reduction, so the reduced o is advanced but ɐ and stressed
ɔ are not: шофьо́р /ʃo̟ˈfʲɔr/.
|
l([^ʲɛi]) → ɫ%1 (repeated) |
ма́лко | ма́лко /ˈmaɫko/ | Includes word-final l (followed by #). |
| l before ɛ, i, ʲ | ле́ка | ле́ка /ˈlɛkɐ/ | Stays l. |
Stage 13: Voicing Assimilation and Nasals
| Rule | Example Input | Output | Description |
|---|---|---|---|
Devoicing: ([bdɡzʒv͡]*)(accent?[ptksʃfx#]) through the devoicing table |
ска́зка | ска́зка /ˈskaskɐ/ | The whole run of voiced obstruents before a voiceless one or # is devoiced:
b→p d→t ɡ→k z→s ʒ→ʃ v→f. Also yields final devoicing: гра́д /ˈɡrat/.
|
Voicing: ([ptksʃfx͡]*)(accent?[bdɡzʒ]) through the voicing table |
про́сба | про́сба /ˈprɔzbɐ/ | p→b t→d k→ɡ s→z ʃ→ʒ x→ɣ f→v. The trigger set does not include v. |
n(accent?[ɡk]+) → ŋ |
ба́нка | ба́нка /ˈbaŋkɐ/ | Velar nasal before к/г. |
m(accent?[fv]+) → ɱ |
ка́мфор | ка́мфор /ˈkaɱfor/ | Labiodental nasal before ф/в. |
| The two substitutions run in this order | се́здесет | се́здесет /ˈsɛzdɛsɛt/ | Devoicing first, then voicing. A commented-out block about clitic с and в (с/в as a prefix of the next word) is present in the source but disabled: в /f/, с /s/ stay as single letters. |
Stage 14-16: Sibilants, Clusters and Cleanup
| Rule | Example Input | Output | Description |
|---|---|---|---|
[sz](accent?[td]?͡?)([ʃʒ]) → %2%1%2 |
разшири́ | разшири́ /rɐʃʃiˈri/ | Sibilant assimilation: the s/z copies the following ʃ/ʒ. The ч case: изчи́стя /iʃˈt͡ʃistʲɐ/; the ж case: изжи́вея /iʒˈʒivɛjɐ/. |
([szʃʒ])[td](accent?)([tdknml]) → %2%1%3 |
че́стно | че́стно /ˈt͡ʃɛsno/ | Deletes т or д between a sibilant and another stop or nasal. Note that the replacement puts the accent before the sibilant. Further examples: по́здно /ˈpɔzno/, ща́стлив /ˈʃtaslif/, по́щна /ˈpɔʃnɐ/. |
# removal |
- | - | All boundary markers are deleted and the string is returned. |
Additional Notes for Developers
- Debug output at load time: the last line of the module before
return exportis a strayprint(export.toIPA("Съединени американски щати")). It runs once when the module is required and writes to the console; it does not affect results. - endschwa is not exposed: the page never passes the second argument. A direct call
toIPA("говоря́т", true)would give the stressed ят as ʲɤ (the dot-below spelling in Common Issues gives the same result with the app). - Tables of unused code:
export.show,export.show_hyphenation,export.remove_pron_notationsandget_anntextexist for the Wiktionary templates{{bg-IPA}}and{{bg-hyph}}and are not called by the app. - Constants:
vowels = "aɤɔuɛiɐo"andcons(including the tie bar) define the character classes used by the stress-movement rules;accentsis the pairˈ ˌ.
Syllabification and Hyphenation (not used by the app)
The same Lua file also contains a sonority-based syllabifier and a hyphenation routine (ported from a project
by "Chernorizets"). The transcription app does not call them, and toIPA does not
depend on them, so they have no effect on the IPA you see. They are exported for the Wiktionary
{{bg-hyph}} template. For completeness:
| Part | What it does |
|---|---|
export.syllabify(term) / syllabify_word |
Splits a word into syllables (joined by ‧ U+2027). A word with one vowel is returned whole. Otherwise, between two vowels: no consonants - break before the second vowel; one consonant - it starts the next syllable; two or more - the break is placed where the sonority stops rising. |
Sonority ranks (get_sonority_rank) |
fricatives 1; stops and affricates 2; sonorants 3; vowels 4. щ counts as ш + т; дж
counts as one stop-rank unit; ь is ignored.
|
sonority_exception_break |
Clusters broken even though sonority rises (км, гм, дм, вм, зм, дн, вн, тн, зд, жд, ... and three-letter ones such as згн, здн, вдж), with the break offset for each. |
sonority_exception_keep |
Clusters kept as the onset: ств, св, вс. A щ + sonorant onset is also
specially handled.
|
prefixes table |
Prefixes (без, из, въз, раз, от, екс, таз, дис, пред and combinations) that are kept as a first syllable when followed by a consonant of higher sonority. |
export.hyphenate(syllabification) |
Applies the official Bulgarian hyphenation rules to a syllabification (never leave a single vowel at a line edge, keep дж together, and so on). |
Examples of the syllabifier output (from export.syllabify, shown as plain text; these are not IPA):
безбрачие → без‧бра‧чи‧е, изкуство → из‧ку‧ство, сладко → слад‧ко. Hyphenation of ябълка gives ябъл‧ка.
The "keep" rule for ств is not available to the IPA stress-movement code, which is why the mark can
fall in a different place in a word like изкуство́ /iskostˈvɔ/.
Common Issues & Limitations
Known Transcription Problems
This table shows known issues where the automatic transcription may be wrong or surprising. Every output below was produced by the real engine.
| Input | System Output | Correct IPA | Cause | What to Do |
|---|---|---|---|---|
| добър (no stress mark) | добър /dobɐr/ | до́бър /ˈdɔbɐr/ | Without a stress mark all vowels are treated as unstressed and reduced, and no ˈ is printed | Type a combining acute after the stressed vowel |
| България (capital, no stress mark) | България /bɐɫɡɐrijɐ/ | Бълга́рия /bɐɫˈɡarijɐ/ | Same cause; capitals are lowercased and do not mark stress | Mark the stress |
| ма̀ма (grave accent only) | ма̀ма /ма̀ма/ | ма́ма /ˈmamɐ/ | The module raises an error for a grave without an acute, and the input comes back unchanged (not IPA) | Use the acute accent U+0301 for the primary stress |
| говоря́т (stressed ending) | говоря́т /ɡovoˈrʲa̟t/ | говоря̣́т /ɡovoˈrʲɤ̟t/ | In 3rd person plural verbs the stressed ending -ят/-ат is pronounced with ɤ; the engine only knows that if you ask for it | Put a dot below (U+0323) under the я or а |
| джоб | джо́б /ˈdʒɔp/ | The same string but with a tie bar, d͡ʒ instead of dʒ | дж and дз are treated as д + ж / д + з, with no affricate marking | Accept, or correct by hand; no input will change it |
| изкуство (stress on the last vowel) | изкуство́ /iskostˈvɔ/ | By the module's own syllabifier the whole onset ств belongs to the stressed syllable, so the mark should precede it; the engine puts it after the t, inside the cluster | The stress-movement rules cover Cl/Cr, kv/gv and s+C but not s+t+v | Treat the position of the mark as approximate in three-consonant clusters |
| ма́ма́ (two acute accents) | ма́ма́ /ˈmaˈma/ | ма́ма /ˈmamɐ/ | No validation: every accent is moved and printed | Use one acute per word (and a grave for secondary stress) |
| До́бър де́н (two words in one string) | до́бър де́н /ˈdɔbɐrˈdɛn/ | Transcribe each word separately | The cleaning step deletes the space, so both words are merged | Enter one word per call |
| чéрен typed with a Latin é (U+00E9) | чéрен /t͡ʃeˈrɛn/ | че́рен /ˈt͡ʃɛrɛn/ | Letters that are not Bulgarian (here Latin e) are passed through unchanged into the IPA | Type the acute as a combining accent U+0301 on a Cyrillic letter; do not use precomposed Latin letters |
| hello (Latin script) | hello /heɫɫo/ | Not a Bulgarian word | No script check; the letters not in the mapping are copied | Check the input is Cyrillic |
Related Pronunciation Guides
Help pages for other Slavic languages in this tool:
For technical issues or suggestions, please visit our GitHub repository.