Downloads
The full corpus, free to reuse
These exports come directly from the project's working database and may occasionally retain internal process annotations (review exchanges, points pending arbitration) within the body of some lexical entries — left as-is for transparency rather than artificially stripped.
Arabic text (Hafs and Warsh readings), segmentation into units of meaning, translations in French, English, Spanish, Italian, Indonesian and Russian, stabilized root lexicon.
Contents : arabic_soura, arabic_phrase, arabic_words, root_knowledge, readings, translations
Format : SQLite (.sqlite) + README + LICENSE · Size : ~10 MB
Analysis infrastructure built on the corpus: root co-occurrences, semantic graph, morphology, grammatical rules derived from the corpus, stabilized thematic clusters.
Contents : root_cooccurrence, semantic_nodes/edges, racine_liens, regles_grammaticales, clusters, root_forme/geste/objet…
Format : SQLite (.sqlite) + README + LICENSE · Size : ~3 MB
Flat, spreadsheet-compatible version of phrases, words and roots — for those who prefer to avoid SQLite.
Contents : phrases.csv, mots.csv, racines.csv, soura.csv
Format : CSV (UTF-8) + README + LICENSE · Size : ~6 MB