Quran content should never become anonymous data. Every Quran.ws dataset identifies its source, edition, riwayah, version, and verification status.
Every dataset starts from an identified printed edition or scholarly reference, never from an unattributed file.
Extraction, normalization, and segmentation steps are listed so the path from source to data is reproducible.
Text and geometry are checked against the source. Status stays "In review" until the audit completes.
Each version ships with a digest. Corrections create a new version; earlier versions remain available.
Select a row to view its full record.
Ḥafṣ text taken unedited from the KFGQPC digital package named above. Word identity is shared with all other riwayat in Quran Text.
Found a discrepancy between a dataset and its printed source? Open an issue with the dataset name, version, digest, and the page or ayah reference. Corrections ship as a new version with a new digest; earlier versions stay available.
Open an issue on GitHub →