A national initiative for Lebanese

Lebanon, in its own words.

A living record of Lebanese — so technology can understand it.

Every day, Lebanese is spoken in ways almost nothing online has recorded: its proverbs, the words of its trades, its mixing of Arabic, French and English, its accents from every town. We're collecting it, with the people who speak it, and keeping it open.

What is Bel-⁠Lebneneh

Bel-⁠Lebneneh is an open initiative to record Lebanese as it is actually spoken, and to keep it usable by the technology of the coming decades.

We're building four things: an open corpus of spoken and written Lebanese; models adapted to it; a community of Lebanese researchers to carry the work; and a way for anyone to contribute their own Lebanese.

Bel-⁠Lebneneh is coordinated by Fondation Diane, a Lebanese nonprofit, in line with Lebanon's national strategy for artificial intelligence. What we collect is a public good.

Why it matters

What our grandparents know is written down nowhere

The generations who grew up before the internet carry three things that exist almost nowhere online. The way people really speak, region by region, with expressions no dictionary holds. The vocabulary of trades — the workshop, the kitchen, the clinic — where Arabic, French and English mix in ways no textbook describes. And local memory: what happened in a village, who lived there, what the place was called before.

When they go, all of this goes with them.

There is a second cost, and it lands on people who are alive now. Speech that machines don't understand ends up shutting out the people who speak it. When technology reads only Modern Standard Arabic, the people furthest from formal schooling are the first to be shut out of the services, the information and the tools everyone else takes for granted.

And in a country where so much is contested, the way Lebanese people speak remains something genuinely shared — across regions, across communities, across generations.

The corpus

Four layers

  1. Spoken

    Conversations, stories and family memory, recorded with consent.

  2. Written

    Lebanese texts from public and community sources.

  3. Digital

    Arabizi, and the mix of Arabic, French and English we use online.

  4. Benchmarks

    Test sets that measure how well a system actually understands Lebanese.

Each layer will be documented to a public standard, so that what we collect can be checked, reused and built on — not just admired.

Bring your Lebanese.

Soon you'll be able to contribute your voice, your proverbs, your family stories and your everyday Lebanese. Every saying you share becomes part of a lasting record of how Lebanon speaks. The first call will be for proverbs and expressions — the sayings you grew up with.

If you have grandparents who tell stories, they are the people we most need to hear.

In Lebanon or abroad, your Lebanese counts.

Your contribution, your rights

  • A receipt for every contribution.

    When you contribute, you download a receipt: a random registration number, the date, and the details you gave us. Keep it — it's how you find your contribution again.

  • You know where it goes.

    Before you record, we explain clearly how your contribution will be used and shared.

  • You can withdraw it.

    Email privacy@bellebneneh.org with your receipt attached, or with your registration number. Before publication, your contribution is deleted entirely. After publication, it is removed from all future releases, but copies already published can't be recalled.

  • Lost your receipt?

    Write to us with the date and place you recorded, and we'll help you find your contribution.

Partners

Bel-⁠Lebneneh is coordinated by Fondation Diane, a Lebanese nonprofit, and is in dialogue with Lebanese universities and institutions.

On Mozilla Common Voice, the open voice platform, Lebanese now has its own variants — one in Arabic script and one in arabizi. Recording opens once the interface is translated and the first sentences are in place.

Fondation Diane

Open by design

Open by default. Nothing without consent.

  • Public domain where it opens doors.

    What we contribute to Mozilla Common Voice is free for anyone, anywhere, to use (CC0).

  • Free for research and teaching.

    The rest of the corpus is free for research and teaching (CC BY-NC-SA). Commercial use needs a separate agreement.

  • Consent first.

    Nothing is recorded without informed consent. Before you record, you know where your contribution goes.

  • Reviewed before release.

    Every contribution is checked before publication. Personal details are never published.

  • Your data, protected.

    We ask only for what the corpus needs. Contributions will open only once our privacy notice is published, setting out how your data is stored and protected.

Stay in touch

News about the project, and first word when contributions open.

Until our mailing list is ready, write to contact@bellebneneh.org and we'll let you know.