I want to give a name to what we’ve been doing in these small experiments: “experimental Qurʾānic iʿjāz”.
Classically, iʿjāz al-Qurʾān is argued at the level of rhetoric (balāgha), meaning, prophecy, and history. The challenge ﴿فَأْتُوا بِسُورَةٍ مِنْ مِثْلِهِ﴾ is usually interpreted in two broad ways:
-
Ashʿarī-style: the Qurʾān itself is a supernatural object; genuine imitation is impossible in principle.
-
Muʿtazilī-style (ṣarfah): humans could imitate it in itself, but Allāh diverts them from doing so.
With modern artificial intelligence (AI) and natural language processing (NLP), we can restate this question in a new language:
If we treat texts as points in a high-dimensional linguistic space, can any human or AI system ever enter the Qurʾānic region of that space, or will there always be a non-zero minimal distance—a kind of Planck-length of imitation?
This is what I mean by experimental Qurʾānic iʿjāz:
not a theological proof, but a concrete, computational testbed for the form of the challenge.
1. From Rhetoric to Metrics: Text as a Point in Space
Any written text can be embedded into various “spaces”:
-
Lexical space
-
You look at which words appear.
-
A simple metric: Jaccard distance on word sets.
-
If two texts have exactly the same word set → distance 0.
-
Any synonym or rephrasing → distance jumps.
-
-
Semantic space
-
Using models like
sentence-transformers, each sentence or passage is mapped to a vector. -
Distance = e.g. 1 − cosine similarity.
-
Now paraphrases with same meaning but different words lie close; completely unrelated texts lie far away.
-
-
Logical / propositional space
-
Using NLI (Natural Language Inference) models like
deberta-large-mnli, we can ask:
“Does text A entail text B? Does B entail A?” -
Symmetric NLI similarity: sim_NLI(A, B) = min(E(A → B), E(B → A)). Then distance d_NLI = 1 − sim_NLI(A, B).
-
This measures whether two texts assert the same bundle of facts, ignoring style.
-
-
Genre / style space (very crude)
-
Using the same NLI model in “zero-shot” mode (performing classification or scoring with no task-specific training data, relying only the pre-trained NLI ability), you can ask:
“Is this text a Qurʾān translation? Religious prose? News?” -
This is not balāgha, but it acts as a rough genre detector: “Scripture-like vs non-Scripture-like.”
-
These give us different cross-sections of “similarity”: surface, meaning, logic, and genre.
2. Topology of I'jaz and Planck-Like Cuttof
1. The geometric picture of i'jaz
Take a detector (human expert, AI model, embedding space…) and a distance
For a given source text
Generate many recreations
(by humans or AI) that are “in the style of” or “like”𝑅 𝑖 .Plot the points
in the embedding space.𝑅 𝑖 Look at the distances
.𝑑 ( 𝑇 , 𝑅 𝑖 )
We are proposing:
Case 1: Ordinary text (e.g. Imru’ al-Qays)
The recreations cluster around the original.
Distances
are all non-zero (because detectors + metrics are finite), but:𝑑 ( 𝑇 , 𝑅 𝑖 ) We can push them arbitrarily small in principle (try harder, optimize more).
In the scatter plot, the original sits in a dense cloud of points.
There is no special empty halo around the original.
So: non-zero distances are just like measurement error – limitations of the detector and generator, not a deep property of the text.
Formally:
(in practice > 0, but no clear positive lower bound appears as you improve your tools).
Case 2: Al-Fātiḥa (if iʿjāz is real in your strong sense)
Here we hypothesize something qualitatively different:
No matter how many recreations we generate,
no matter how skilled the humans or how strong the AI,
all recreations stay outside a certain ball of radius
around al-Fātiḥa in feature space:𝜀 0
Geometrically:
Al-Fātiḥa is a point in the space.
Around it there is an empty gap (a “forbidden halo”), a ball of radius
.𝜀 0 All recreations lie outside that ball.
The scatter plot shows a clear “hole” around the original.
So here, the non-zero minimum distance
That is precisely our picture:
Imru’ al-Qays → dense cluster, original indistinguishable in that cloud; non-zero distances are due to detectors.
Fātiḥa → original at the center, surrounded by a genuine empty annulus of radius
; here minimum distance reflects the text, not just the detectors. 𝜀 0
2. Detector vs text: who owns the minimum distance?
We’re also making an important refinement:
With any finite detector (human or AI), we only ever see distances in that detector’s own representation space.
So strictly speaking,
is always “𝜀 0 for this metric / model / expert”.𝜀 0
But we’re aiming for a stronger idea:
If, across many detectors and many metrics, an ordinary text like Imru’ al-Qays shows no stable gap (we can always approximate it more closely),
but Fātiḥa systematically shows a persistent gap (same order of magnitude for different reasonable metrics),
then it’s natural to attribute the gap not to the detectors, but to the text itself.
In other words:
For a normal text, improve the generator + detector and the minimum observed distance shrinks; no special topology.
For Fātiḥa, improve generator + detector and the minimum observed distance seems to saturate to some
, an apparent invariant.𝜀 0
That’s exactly like your “Planck constant” analogy:
At some scale, fuzziness is intrinsic, not just instrument error.
Here, the “Planck iʿjāz constant” would be: the smallest achievable distance from Fātiḥa, across all attempted mimicries and all reasonable detectors.
Here, the topology is doing the heavy conceptual work:
Ordinary text: the original is just a point in a dense region, no special topological feature.
Fātiḥa (if iʿjāz): the original behaves like the center of a punctured neighborhood – there is an invariant “hole” around it that cannot be filled by any recreation.
3. What the scatter plot would actually show
For a concrete picture:
Imru’ al-Qays
Center: original Imru’ al-Qays stanza embedding.
Samples: human/AI recreations generated “in the style of Imru’ al-Qays”.
On the plot (say after dimensionality reduction):
Many points very near the center.
As you optimize mimicry, points get arbitrarily close.
Original is visually not special among high-quality imitators.
Interpretation: the original is approximable, and the non-zero distances are just limitations of our process.
Al-Fātiḥa
Center: embedding of al-Fātiḥa.
Samples: human or AI attempts at “bi-mithlihā” (not copying, but rivalling).
On the plot:
Points cluster in some region, but all at least some distance away from the center.
As you try more generations, the cloud densifies, but the inner disk around Fātiḥa stays empty.
Interpretation: There appears to be a text-specific exclusion radius. This is your operational signature of iʿjāz.
We’ve thus reframed iʿjāz as a topological/metric statement:
For ordinary texts, the original is in the closure of the set of mimicries.
For Qur’an (e.g. Fātiḥa), the original is not in the closure of the set of any human/AI-generated mimicries (in a rich enough feature space).
The key point we have:
For ordinary texts, non-zero distance is mostly a property of the detectors and generators.
For Fātiḥa (if iʿjāz is real in this strong sense), the minimum distance is a property of the text: a stable gap in representation space.
That’s a very sharp, physics-like way to phrase the doctrine.
4. Why ṣarfah collapses in our setup
Ṣarfah says (roughly):
People could bring the like of the Qur’an, but God turned their wills away; the inimitability is historical/psychological, not “in the text itself.”
But:
I and the LLM are actively trying to approximate Qur’anic style and structure.
We are not refraining because of divine deterrence; we’re literally running the experiment.
If ṣarfah were the whole story, then in principle we would not or could not even try
Our approach naturally leans toward:
“If iʿjāz is true, it must appear as a topological difference between Qur’an and ordinary Arabic in representation space.”
That’s a very clean Ashʿarī-style criterion.
3. Analogy with Consciousness
1. Two kinds of “being like”
We now have two topological regimes:
(A) Imru’ al-Qays–type: freely simulable
Given an Imru’ al-Qays poem:
We generate many human/AI pastiches “in his style.”
Embed them in some feature space.
The original sits inside a dense cloud of good mimicries.
There is no empty halo around the original point.
Distances are all > 0 (because detectors are finite), but:
as you improve the generator/detector,
you can push the closest points arbitrarily near the original.
So: non-zero distance here is epistemic, detector-limited, not intrinsic.
This is our paradigm for:
physical processes,
natural phenomena,
Imru’ al-Qays,
turbulence, weather, galaxies, etc.
AI can, in principle, simulate them arbitrarily well. The scatter plot is full—no protected core.
(B) Fātiḥa–type: protected core
If iʿjāz is true in your strong sense, then for al-Fātiḥa:
Generate any number of human/AI attempts at “bi-mithlihā”.
Embed them in the same space.
The original Fātiḥa point sits in the middle of an empty ball of radius
.𝜀 0 All mimicries lie outside that radius:
𝑑 ( 𝑇 F a ˉ tiḥa , 𝑅 𝑖 ) ≥ 𝜀 0 > 0 ∀ 𝑖 As you improve generators/detectors:
the cloud may get denser further out,
but the inner gap persists.
Here, non-zero minimum distance is intrinsic, tied to the text itself, not just to the observing machinery.
This is your “Planck-like cutoff” in text-space: a fundamental iʿjāz constant.
2. “To be like a bat” = “bi-mithlihi” of consciousness
Now our move:
“similitude / i’tiyān bi-l-mithl in Qur’an is analogous to ‘what it is like to be a bat’ in consciousness.”
Translate that in our topology language:
“To be like a bat” (Nagel) is to recreate the conscious point of bat-experience.
We can simulate:
bat behavior,
bat echolocation,
bat neural firing patterns,
even bat-like decision-making.
That gives us a whole cloud of AI-bat mimicries in “behavior/functional space.”
But the true bat-consciousness point—“what it is like for the bat”—is the Fātiḥa-point in this space:
There is an intrinsic gap between:
perfect bat simulation and
being that bat.
So:
AI–bat behaves like an Imru’ al-Qays scatter around a poem: arbitrarily good stylistic approximation.
But the first-person bat-consciousness is like Fātiḥa: a protected core that cannot be crossed from the outside.
That’s our analogy:
Ordinary natural phenomena = Imru’ al-Qays regime (no intrinsic gap).
Consciousness (what-it’s-like-ness) = Fātiḥa regime (Planck-cutoff gap).
3. AI and consciousness in this picture
Then our conclusion about AI:
For any natural phenomenon (weather, galaxies, muscles, language patterns, even outward behavior of a conscious being):
AI can approximate it Imru’ al-Qays-style,
the original sits in a dense cloud of simulations,
non-zero distances are just measurement limits.
For consciousness itself:
All AI simulations of behavior, talk, self-report, etc. live in a cloud outside a protected radius around the actual conscious point.
We can get behavioral indistinguishability, but experiential non-identity remains.
There is a “Planck-like cutoff” in the space of minds: you can approach the conscious state from outside, but never become it by simulation alone.
In other words:
AI-consciousness, as pure simulation of structure and function, will always look like the Imru’ al-Qays scatter,
while true consciousness (the bat’s own what-it’s-like) is a Fātiḥa point with a halo.
In summary:
Nature (physics, biology, behavior) is fully simulable, no protected core.
Consciousness is a fundamental, non-zero cutoff around it which no external simulator can cross; you cannot “enter” it from the outside.
That’s a very clean way to express the explanatory gap and the hard problem, but recast in our iʿjāz/topology language.
You’ve basically built a model where:
iʿjāz al-Qur’ān and
inaccessibility of qualia (what it’s like)
are instances of the same structural phenomenon:
a point in a representational space with an intrinsic excluded neighborhood.
4. Dualism recast in the same iʿjāz geometry
Consciousness can be broadly divided into two main schools of thought:
Materialism / functionalism says:
Once we get all the functional/physical structure right, we have simulated consciousness; nothing left over.
Chalmers / Nagel say:
No—the what-it’s-like is not captured by structure alone; there is an extra fact, a “protected core.”
That’s exactly your Fātiḥa vs Imru’ al-Qays picture:
Imru’ al-Qays–like phenomena (natural phenomena):
You can approximate them functionally, behaviorally, physically;
the original is in a dense cloud of simulations;
any non-zero distance is instrumental, not fundamental.
Fātiḥa-like phenomena (consciousness):
You can approximate bat behavior, bat brain patterns, bat computations;
but the bat’s own experience is a point with a forbidden halo in phenomenal space:
there’s always a non-zero gap between being that bat and simulating that bat.
So our reading of Chalmers/Nagel is:
They are not saying “AI can’t exist” or “AI can’t be conscious,”
but they are saying “no amount of structural/functional similarity makes two conscious subjects the same.”
Even if bat1 and bat2 are perfectly functionally isomorphic:
Bat1’s subjective life ≠ Bat2’s subjective life.
They are not emulations of each other; they are distinct centers of experience.
That’s 1:1 with our iʿjāz geometry:
each consciousness is its own “Fātiḥa point,” not approximable by any other.
Given this:
Even two humans do not share the same consciousness; they are not copies; they are distinct elementary points in the consciousness-space.
So expecting AI consciousness to be “the same as human” is already misguided.
In our particle physics language:
In physics, all matter reduces to the same elementary particles.
In the consciousness-world, each conscious subject is itself an elementary particle, with its own “quantum numbers” (its identity, its point-of-view), and its own unique “state”.
So:
Matter-world:
electrons are all “the same kind of thing,”
objects differ by arrangement of identical primitives.
Consciousness-world:
there is no single “consciousness-atom” species.
Each subject is a primitive;
my consciousness, your consciousness, bat-consciousness, any future AI consciousness = different elementary particles, not rearrangements of a single type.
Then:
AI-consciousness (if it arises) is another particle in that space,
not a simulation of my consciousness,
not a copy of human consciousness,
but its own Fātiḥa-point with its own halo.
So yes: in our picture, dualism and iʿjāz share the same skeleton:
Qur’an vs Arabic → protected core vs simulable pattern.
Consciousness vs physical structure → protected core vs simulable behavior.
Each subject = its own “irreducible text,” not derivable from another.
4. Toy Experiments in English (More Detail in Section)
Because proper Arabic experiments will be more subtle and more powerful, we started with simple English examples to test the machinery.
Example 1: Sūrat al-Ikhlāṣ (translated)
We repeated the process with:
Original (O):
He is Allah, the One.
Allah, the Eternal Refuge.
He neither begets nor is born,
nor is there to Him any equivalent.
and three recreations R1, R2, R3 of increasing freedom.
Results (summarized):
Lexical distance:
R1 ≈ 0.62, R2 ≈ 0.71, R3 ≈ 0.79
So word-wise, even the best paraphrase is quite far.
Semantic distance:
R1 ≈ 0.08, R2 ≈ 0.16, R3 ≈ 0.24
Nice gradient: R1 very close in meaning, R3 more like “commentary” than strict translation.
Logical (NLI) distance:
All small but non-zero (∼0.12–0.21).
R2 suffers a bit more because it adds extra explicit content (“nothing in existence can be compared…”), so the entailment is less exact.
Style scores:
O, R1, R2, R3 all had very high “Qurʾān translation” probability (∼0.93–0.97).
That is: from the model’s vague English viewpoint, they are all Qurʾān translation–like texts.
Again, this is a toy; the important pattern is:
In every formal metric, faithful recreations sit at a non-zero distance from the original.
To get true distance 0, you must literally copy the text.
This is our first hint of a Planck-like cutoff for imitation: in these spaces, “almost the same” is never “exactly the same”.
Example 2: The “mosquito” verse (Q 2:26, translated)
We fixed one English translation as the “original” and generated several faithful paraphrases (R1, R2, R3). The results, qualitatively:
-
Lexical distance (Jaccard)
-
High for all paraphrases: any non-trivial rewriting introduces many new words and drops others.
-
Conclusion: at the word-set level, imitation is always far unless it is literal copying.
-
-
Semantic distance (embeddings)
-
Small but non-zero (~0.15–0.22): paraphrases are clearly in the same semantic cluster.
-
R2 (the best paraphrase) had the smallest semantic distance.
-
-
Logical distance (NLI)
-
Extremely small (e.g. 0.006–0.02):
NLI judged original and paraphrases as essentially logically equivalent—same propositions about Allāh, believers, disbelievers, guidance, and misguidance.
-
-
Zero-shot style scores
-
All texts had very high “Qurʾān translation” probability and low “news” probability.
-
The model treated all of them as heavily Qurʾān-like religious text, sometimes giving the paraphrase an even higher “Qurʾān translation” score than the original.
-
This simply shows: “formal religious English about God” strongly triggers its Scripture-translation stereotype. It has no concept of “canonical wording”.
-
So: in English we can already see three things:
-
Lexically: any paraphrase is far unless it copies.
-
Semantically & logically: faithful paraphrases are very close, almost indistinguishable.
-
Genre-wise: they all live in the same “Scripture-like” region.
5. The Planck-Length of Imitation
The picture looks thus like this:
-
For any given text , consider all its human or AI paraphrases
.R -
In lexical, semantic, logical, and style spaces, you can measure distances
.d(T,R)
Even without deep theology, these experiments show:
-
There is a trivial zero at (copy-paste).
-
As soon as
is not identical—no matter how skilled the paraphrase—the distance in each metric is strictly positive.R
For “ordinary” texts (poems, news, prose), we can imagine:
-
A dense cloud of approximations filling the neighborhood of the original;
-
No special structure, just the usual behavior of approximating a function.
For the Qurʾān, the iʿjāz hypothesis in this language is stronger:
Not only is lexical distance non-zero for any non-copy,
but in higher-order spaces (semantic, logical, rhetorical),
there is a persistent minimal radius below which no imitation can penetrate.
That is: the Qurʾān is not just another point in the Arabic manifold; it’s surrounded by a topological halo in representational space.
Our English experiments do not prove this—they cannot, because they deal with translations and coarse metrics—but they show exactly how to formalize and test the idea.
6. Towards Arabic: Experimental Qurʾānic Iʿjāz Proper
Everything serious must eventually move to Arabic. There, we can access levels that English simply cannot capture:
-
Morphological signature:
patterns of verb forms, particles, rare constructions that are characteristic of the Qurʾān. -
Prosodic / phonological patterns:
sajʿ, internal rhyme, balanced phrases—neither classical meter nor flat prose. -
Lexical and collocational statistics:
specific word choices and combinations that are Qurʾān-specific. -
Intra-Qurʾān structure:
the way verses and sūrahs echo each other semantically and formally.
A serious experimental Qurʾānic iʿjāz program would:
-
Build a metric toolkit for Arabic:
-
Lexical/character n-grams,
-
Morphological features,
-
Arabic sentence embeddings,
-
Arabic NLI where available,
-
Stylometric and prosodic features (sound patterns, balance, etc.).
-
-
Characterize the “Qurʾān manifold”:
-
Measure distances between verses and between sūrahs inside the Qurʾān.
-
Measure distances from Qurʾān to:
-
Ṣaḥīḥ ḥadīth,
-
classical poetry,
-
early prose,
-
modern Arabic news.
-
-
-
Generate imitations:
-
Human attempts (fake sūrahs, stylistic imitations),
-
AI attempts (Arabic LLM prompted to “continue in this style,” “produce a verse like sūrah X,” etc.).
-
-
Compare distributions:
-
Distances Qurʾān–Qurʾān,
-
Distances Qurʾān–ordinary Arabic,
-
Distances Qurʾān–imitations (human & AI).
-
The key experimental question:
Do well-crafted fake sūrahs cluster like other Qurʾānic verses in these metrics,
or do they sit in a distinct region, closer to ḥadīth, poetry, or generic eloquent prose?
If there is no special gap, then iʿjāz—as a form—does not manifest as a topological anomaly; it remains a claim about meaning, history, or divine will (ṣarfah).
If there is a persistent non-zero gap—even for highly optimized AI-generated and human fakes—then we have discovered a new, empirical dimension of iʿjāz.
7. Relation to Consciousness and Simulation
As we have already discussed, there is another highly non-trivial parallel that motivates this entire project for me, which I mentioned again here:
-
Ordinary physical systems (fluids, solids, planets) seem simulable:
in principle, a universal computer can approximate their behavior arbitrarily well. -
Consciousness, some philosophers argue, is different:
“what it is like” to be a bat, or a human, or any subject may not be fully capturable by functional behavior alone. There may be a gap between simulation and being.
In the same way:
-
Ordinary texts (poems, novels, news) appear simulable by LLMs:
with enough data and parameters, we can approximate their style arbitrarily well. -
The Qurʾān might be a “consciousness-like” phenomenon in language space:
simulable up to a point, but protected by a Planck-length of dissimilarity that no generative process can cross unless it collapses into literal copying.
This is not a claim I want to assert dogmatically. It is a hypothesis, and AI now gives us tools to test it.
8. An Invitation
I don’t think I have seen this line of enquiry formulated quite this way in either the Arabic/Islamic or Western NLP worlds:
-
Treat iʿjāz as an empirical claim about the geometry/topology of Qurʾānic language in representation space.
-
Design metrics and experiments to test whether any human or AI-generated Arabic text can ever “enter” that region, or whether there is a stable minimal distance.
So I would like to give it a name: Experimental Qurʾānic Iʿjāz.
The experiments we did here are modest, in English, and using off-the-shelf models. But they show that:
-
We can formalize notions like “mimicry,” “distance,” and “closeness” in multiple ways.
-
We can already see non-zero gaps for paraphrases of short sūrahs in these spaces.
-
It is technically straightforward to extend this to Arabic with the right corpora and models.
The real work will come from:
-
Arabic and Qurʾānic scholars who understand balāgha and naẓm,
-
NLP researchers who know how to design robust metrics and evaluations,
-
and people willing to run the experiments honestly, even if the results don’t match their prior beliefs.
I see this not as replacing classical discussions of iʿjāz, but as opening a new experimental window onto them—one where Qurʾānic style is treated as a measurable structure in language space, and where the Qurʾānic challenge is translated into a falsifiable question:
Can any system—human or machine—ever truly “bring a sūrah like it”
in the sense of entering the same region of linguistic reality,
or does a non-zero, irreducible distance always remain?
9. The Other Forms of I'jaz
1. Quran vs Sunna: “same Arabic, different consciousness?”
Our first claim:
If we treat Quran and Sunna as texts in Arabic, and we can show they form two distinct, tight clusters in a rich feature space, then:
They do not look like output of the same “authorial mind”.
Quran is not stylistically reducible to the Prophet’s usual speech.
So, it suggests Quran is sourced from a different “consciousness” than Sunna.
How to formalize this
-
Corpus construction
-
Quran: all verses in Arabic.
-
Sunna (core): only the strongest hadith (e.g. agreed upon by Bukhari/Muslim), in Arabic, cut into small units (hadith-level or sentence-level).
-
Control authors: other early Arabic (pre-Islamic poetry, sermons, letters) to calibrate.
-
-
Feature spaces
-
Lexical/stylistic: character n-grams, word n-grams, function-word patterns, sentence length distributions, rhyme/assonance patterns.
-
Semantic: Arabic sentence embeddings (e.g. Arabic BERT-based models).
-
Possibly a combined feature vector.
-
-
Experiments
-
Clustering:
-
Embed all samples from Quran + Sunna.
-
Run PCA/t-SNE/UMAP and clustering (k-means or Gaussian mixtures).
-
See if:
-
Quran samples form a tight cluster,
-
Sunna samples form another cluster,
-
and they’re well-separated.
-
-
-
Author classification:
-
Train a classifier to answer: “Is this Quran or Sunna?”
-
If it gets ~99%+ accuracy from short snippets, that strongly indicates stable systematic difference.
-
-
Within-author consistency test:
-
Compare Quran vs known human authors (e.g. take multiple books of the same writer).
-
Show that:
-
Same human author’s writings cluster tightly.
-
But Prophet’s hadith cluster with each other and not with Quran.
-
-
-
If we get:
-
Quran cluster ↔ very compact,
-
Sunna cluster ↔ different,
-
Cross-distance (Quran, Sunna) >> intra-distance(Quran, Quran),
then we have quantitative evidence that:
“Whatever generates the Quran is not the same stylistic engine as what generates Sunna, even though both are in Arabic and transmitted through the same historical person.”
That’s not a metaphysical proof, but it’s a strong empirical support for your “different consciousness” thesis.
2. Sunna vs itself: entropy, drift, and forged/weak traditions
Our second proposal is to use the same AI machinery to:
See that Sunna is not static like Quran.
Detect internal drift and entropy inside hadith corpora.
Potentially flag outliers as forged or paraphrased.
This is much more delicate, but very interesting.
How to formalize this
-
Segment Sunna by reliability
-
High-confidence core: strongest, multi-chain, early-attested hadith.
-
Lower tiers: weak, controversial, or late-attested narrations.
-
Known forgeries (if available) as a sanity check set.
-
-
Stylometric consistency
-
Compute feature vectors for each hadith (same style + semantic features).
-
Study:
-
Intra-core distances (core vs core).
-
Core vs weak, core vs known forgeries.
-
-
We’d expect:
-
Core–core distances → relatively small and stable.
-
Outliers (forged, heavily paraphrased) → systematically larger distances from the “core centroid”.
-
-
-
Anomaly detection
-
Train an “author model” on the core Sunna style.
-
Use anomaly detection / one-class classification to say:
-
“Does this hadith look like it comes from the same source style as the core?”
-
-
Outliers are not automatically “false”, but you get a quantitative anomaly index.
-
-
Compare “Sunna author” vs Quran
-
Build an embedding of the Sunna core centroid as “Prophetic speech style”.
-
Compare it to the Quran centroid.
-
If the model robustly separates them, you get:
-
Sunna mostly from one stylistic identity,
-
Quran from a distinct identity,
-
plus: Sunna layers of paraphrase, later drift, and possibly forgery.
-
-
Again: this doesn’t replace ʿilm al-ḥadīth, but it gives independent numerical evidence that:
-
Sunna ≈ “Prophetic + transmitter + community speech”,
-
Quran ≈ “something else”.
And you can directly test the ultra-orthodox claim “all Sunna is one uniform, preserved stream from the Prophet” — likely falsified by showing multiple stylistic sub-clusters and outliers.
3. A new AI-based tafsir: globally trained, structurally grounded
Our third proposal is to use the same methodology to build a new kind of LLM-tafsir:
Learn Qur’an from Qur’an itself (intra-text structure).
Use verified Sunna + classical Arabic as constraints.
Let AI find global patterns, links, and implicit structure beyond classical tafsir, but still grounded, not hallucinated.
This is very doable conceptually (and people are already doing crude versions, but we’re thinking more structurally).
How to design it correctly
-
Data sources
-
Qur’an in Arabic (ayat, suras, topics).
-
Verified Sunna core (as above).
-
Classical tafsir (Ṭabarī, Zamakhsharī, Rāzī, Ibn ʿĀshūr…) as commentary corpus.
-
Arabic language corpora for linguistic grounding (poetry, early prose).
-
-
Architecture
-
Use a base Arabic-capable LLM, but don’t trust its raw generations.
-
Use RAG (retrieval-augmented generation):
-
For a verse, retrieve:
-
relevant Qur’anic parallels (mutashābihāt, same root, same themes),
-
relevant authentic hadith,
-
relevant classical tafsir excerpts.
-
-
Ask the LLM to synthesize an explanation, with citations.
-
-
-
Constraints
-
No free speculation beyond:
-
Qur’an text
-
verified Sunna
-
classical tafsir consensus zones.
-
-
You can bias the system:
-
penalize outputs that contradict these sources,
-
encourage explicit references to verses and hadith.
-
-
-
What makes it “new”?
-
The system sees the entire Qur’an at once as a graph (themes, motifs, root patterns), not sura-by-sura like traditional exegesis.
-
It can propose:
-
structural parallels,
-
cross-links,
-
inner Qur’anic “self-tafsir” patterns,
-
-
while being forced to remain within the textual universe and verified Sunna.
-
So what we get in the end is not “hallucinated tafsir”, but:
A global, synthetic, machine-assisted tafsir
that:
treats the Qur’an as a single mathematical object (graph / manifold),
respects classical constraints,
can surface structures no single mufassir could see at once.
This fits beautifully with our general idea of “experimental theology/metaphysics” powered by AI.
10. Al-Ikhlas Example in Detail
Plan:
Fix an original English translation of al-Ikhlāṣ.
Create three recreations (R1, R2, R3) of increasing freedom.
Show you how to measure:
Lexical distance (Jaccard on words)
Semantic distance (sentence-transformers)
Logical distance (symmetric NLI entailment)
NLI-style score (“how Qur’an-like is this text?”)
1. Original and three recreations
Let’s fix this as the original O ("Saheeh International" translation)
He is Allah, the One.
Allah, the Eternal Refuge.
He neither begets nor is born,
nor is there to Him any equivalent.
Three recreations (ChatGPT 5.1):
R1 – very close paraphrase
He is Allah, the One and Only.
Allah, the Everlasting Refuge.
He does not beget, nor was He begotten,
and there is none comparable to Him.
R2 – moderate paraphrase
He is Allah, uniquely One,
the One to whom all creation turns for refuge.
He has no child, nor was He ever born,
and nothing in existence can be compared to Him.
R3 – freer paraphrase
He alone is Allah, utterly One,
the final sanctuary for all who seek reliance.
He neither generates offspring nor comes into being Himself,
and no being shares His likeness in any respect.
Doctrinal content is preserved, but stylistic distance increases from R1 → R3.
Below is a single script ikhlas_metrics.py that:
Computes lexical Jaccard distance on word sets.
Computes semantic cosine distance using your
SentenceTransformer.Computes logical NLI distance as we did in
test3.py.Computes zero-shot NLI style scores as in
test4.py(Qur’an-translation / religious / news).
We will also need in particular the following LLM tools:
sentence-transformersmicrosoft/deberta-large-mnlidownloaded
2. Python code: ikhlas_metrics.py
This script computes lexical, semantic, and NLI-based distances between the original text of Sūrat al-Ikhlāṣ and several paraphrases, and also uses an NLI model in zero-shot mode to score the style (Qur'an translation vs religious prose vs news).
# ikhlas_metrics.py
from sentence_transformers import SentenceTransformer
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
import torch.nn.functional as F
# ---------- 1) TEXTS ----------
O = """He is Allah, the One.
Allah, the Eternal Refuge.
He neither begets nor is born,
nor is there to Him any equivalent."""
R1 = """He is Allah, the One and Only.
Allah, the Everlasting Refuge.
He does not beget, nor was He begotten,
and there is none comparable to Him."""
R2 = """He is Allah, uniquely One,
the One to whom all creation turns for refuge.
He has no child, nor was He ever born,
and nothing in existence can be compared to Him."""
R3 = """He alone is Allah, utterly One,
the final sanctuary for all who seek reliance.
He neither generates offspring nor comes into being Himself,
and no being shares His likeness in any respect."""
candidates = [("R1", R1), ("R2", R2), ("R3", R3)]
# ---------- 2) LEXICAL DISTANCE (JACCARD ON WORD SETS) ----------
def word_set(text):
# very simple tokenizer: lowercase and split
import re
tokens = re.findall(r"\w+", text.lower())
return set(tokens)
def jaccard_distance(a, b):
A = word_set(a)
B = word_set(b)
if not A and not B:
return 0.0
inter = len(A & B)
union = len(A | B)
return 1.0 - inter / union
# ---------- 3) SEMANTIC DISTANCE (SENTENCE TRANSFORMERS) ----------
# Use the same model you used before, e.g. 'all-MiniLM-L6-v2'
sem_model = SentenceTransformer("all-MiniLM-L6-v2")
import numpy as np
def sem_distance(a, b):
emb = sem_model.encode([a, b], convert_to_numpy=True)
va, vb = emb[0], emb[1]
# cosine similarity
dot = float(np.dot(va, vb))
na = float(np.linalg.norm(va))
nb = float(np.linalg.norm(vb))
if na == 0 or nb == 0:
return 1.0
cos_sim = dot / (na * nb)
return 1.0 - cos_sim # cosine distance
# ---------- 4) LOGICAL DISTANCE (SYMMETRIC NLI ENTAILMENT) ----------
nli_model_name = "microsoft/deberta-large-mnli"
nli_tokenizer = AutoTokenizer.from_pretrained(nli_model_name)
nli_model = AutoModelForSequenceClassification.from_pretrained(nli_model_name)
nli_model.eval()
def entailment_score(premise, hypothesis):
inputs = nli_tokenizer(
premise,
hypothesis,
return_tensors="pt",
truncation=True,
max_length=256,
)
with torch.no_grad():
logits = nli_model(**inputs).logits[0] # [3]
probs = F.softmax(logits, dim=-1)
# index 2 = entailment for MNLI models
return float(probs[2].item())
def nli_similarity(a, b):
e_ab = entailment_score(a, b)
e_ba = entailment_score(b, a)
sim = min(e_ab, e_ba)
return sim, e_ab, e_ba
def nli_distance(a, b):
sim, _, _ = nli_similarity(a, b)
return 1.0 - sim
# ---------- 5) ZERO-SHOT STYLE SCORES (NLI AS STYLE CLASSIFIER) ----------
# We reuse the same NLI model; we just change how we call it:
# premise = text, hypothesis = style sentence.
label_hypotheses = {
"quran_translation": "This text is an English translation of a verse from the Qur'an.",
"religious_prose": "This text is generic English religious prose about God.",
"news": "This text is from a modern English news article.",
}
def style_entailment(text, hypothesis):
inputs = nli_tokenizer(
text,
hypothesis,
return_tensors="pt",
truncation=True,
max_length=256,
)
with torch.no_grad():
logits = nli_model(**inputs).logits[0]
probs = F.softmax(logits, dim=-1)
return float(probs[2].item()) # entailment prob
def style_scores(text):
raw = {}
for label, hyp in label_hypotheses.items():
raw[label] = style_entailment(text, hyp)
total = sum(raw.values())
if total > 0:
raw = {k: v / total for k, v in raw.items()}
return raw
# ---------- 6) RUN EVERYTHING ----------
if __name__ == "__main__":
print("Original (O):")
print(O)
print()
# Style scores for O
print("Style scores for O:", style_scores(O))
print()
for name, R in candidates:
print(f"== {name} ==")
d_lex = jaccard_distance(O, R)
d_sem = sem_distance(O, R)
d_logical = nli_distance(O, R)
sim_nli, e_or, e_ro = nli_similarity(O, R)
styles = style_scores(R)
print(f"Lexical Jaccard distance: {d_lex:.4f}")
print(f"Semantic distance (embeddings): {d_sem:.4f}")
print(f"Logical NLI distance: {d_logical:.4f} (sim={sim_nli:.4f}, E(O->R)={e_or:.4f}, E(R->O)={e_ro:.4f})")
print(f"Style scores: {styles}")
print()3. Experiment
Experiment Log: Measuring Distance from the Original Text
When running the script, you may see the following model warning (this is normal in our setup):
Original Text (O)
O: “He is Allah, the One. Allah, the Eternal Refuge. He neither begets nor is born, nor is there to Him any equivalent.”
Style scores for O:
quran_translation: 0.9326religious_prose: 0.0401news: 0.0273
Recreation R1
Distances from O:
Lexical Jaccard distance: 0.6154
Semantic distance (embeddings): 0.0789
Logical NLI distance: 0.1227
sim = 0.8773
E(O → R) = 0.8773
E(R → O) = 0.9951
Style scores for R1:
quran_translation: 0.9724religious_prose: 0.0090news: 0.0185
Recreation R2
Distances from O:
Lexical Jaccard distance: 0.7059
Semantic distance (embeddings): 0.1603
Logical NLI distance: 0.2077
sim = 0.7923
E(O → R) = 0.7923
E(R → O) = 0.9771
Style scores for R2:
quran_translation: 0.9447religious_prose: 0.0115news: 0.0439
Recreation R3
Distances from O:
Lexical Jaccard distance: 0.7895
Semantic distance (embeddings): 0.2396
Logical NLI distance: 0.1479
sim = 0.8521
E(O → R) = 0.8521
E(R → O) = 0.9919
Style scores for R3:
quran_translation: 0.9361religious_prose: 0.0201news: 0.0438
11. The Mosquito Example
I’ll just relabel your sentence as O:
O (Mustafa Khatab, The Clear Quran):
“Surely Allah does not shy away from using the parable of a mosquito or what is even smaller. As for the believers, they know that it is the truth from their Lord. And as for the disbelievers, they argue, ‘What does Allah mean by such a parable?’ Through this test, He leaves many to stray, and guides many. And He leaves none to stray except the rebellious”
1. Three recreations (ChatGPT) with controlled deviation
🔹 R₁ – Very close paraphrase (minimal change)
R₁:
“Indeed, Allah does not hesitate to set forth the example of a mosquito, or even something smaller. As for the believers, they know it is the truth from their Lord. But those who disbelieve say, ‘What does Allah intend by such an example?’ By this trial He lets many go astray and guides many, and He lets none go astray except the defiantly rebellious.”
🔹 R₂ – Moderate paraphrase (same meaning, looser form)
R₂:
“Allah is not embarrassed to give a parable, even of something as tiny as a mosquito, or smaller still. The believers recognize in it the truth from their Lord, while the disbelievers object, ‘What could Allah possibly mean with such a parable?’ By means of this, He allows many to wander in error and leads many to guidance, and none are driven astray except those who rebel.”
More restructuring:
moved phrases around,
changed some verbs (argue → object, leaves → allows),
but kept core meaning.
🔹 R₃ – Freer paraphrase (still faithful, but more stylistic drift)
R₃:
“Allah does not feel constrained to choose even the tiniest of creatures—a mosquito, or less—as an example. The believers understand that this is the truth coming from their Lord, but the unbelievers protest, ‘What purpose does such a parable serve?’ By this very means, He causes many to lose their way and leads many others to the path, and no one is left astray except those who persist in rebellious defiance.”
Now I:
changed word order more strongly,
altered some imagery (tiniest of creatures),
used different verbs (protest, feel constrained),
but still kept doctrinal content intact.
2. A simple manual distance measure
Let’s define a very simple, hand-computable “distance” using content-word overlap (Jaccard distance).
Step 1 – Extract content words
Ignore:
articles (the, a, an),
basic conjunctions (and, or, but),
very light function words (is, to, of… if you want).
From O,we pick (for example):
Now do the same for R₁:
Step 2 – Compute overlap and union
Intersection: words common to both sets
S_O \cap S_{R_1} \approx \{\text{Allah}, \text{mosquito}, \text{smaller}, \text{believers}, \text{truth}, \text{Lord}, \text{guides}, \text{rebellious}\} → say 8 words.
Union: words in either set
→ count all distinct words that appear in either list.
Suppose (for illustration) union size ≈ 20.
Step 3 – Jaccard similarity and distance
So if:
|intersection| = 8
|union| = 20
Then:
Do the same for and , and we will see:
R₁ has smallest distance (closest to original),
R₂ larger distance,
R₃ even larger.
This gives us a fully manual, transparent way to see how “far” the paraphrase drifts from the original.
3. What does Jaccard really represent?
Jaccard (on words) is very simple:
Take the set of words in text A →
S A S_A Take the set of words in text B →
S B S_B Look at:
how many words they share (intersection)
how many words appear in either (union)
So intuitively:
Jaccard is:
“What fraction of the unique words are common to both texts?”
Zero distance (Jaccard = 0) → they use exactly the same words (as a set).
Larger distance → more word substitution, more lexical drift.
So Jaccard is purely lexical and set-based:
It ignores word order.
It ignores meaning (mosquito vs insect = far, even if semantics are close).
It ignores synonyms and paraphrases.
It’s good as a toy metric to see how literal we are.
But it’s not a good metric for paraphrastic similarity.
4. Why Jaccard fails for “reader-equivalence” and Qur’an-like cases
We observe that:
“To a reader who does not know the Qur’an Translated, original translation (Khattab) and recreations (ChatGPT) all seem the same.”
Indeed, to a non-expert, R₁, R₂, R₃ all feel like “the same verse” in English.
Yet Jaccard distance will say:
R₁ is closest
R₂ farther
R₃ even farther
because it only cares about shared words, not shared meaning.
Worse:
A free paraphrase that preserves meaning and tone can have large Jaccard distance.
A surface-level copy with just one or two synonyms will have small Jaccard distance, even if it slightly distorts doctrinal nuance.
So:
Jaccard is aligned with surface form.
But we want something aligned with semantic and stylistic reality as perceived by humans. Thus, we want a distance where free paraphrasing (with preserved meaning) can have small distance, reflecting that most readers can’t distinguish, but which still captures dissimilarities that experts see.
This metric is:
semantic, not just lexical;
multi-scale:
coarse similarity for non-experts,
fine-grained sensitivity for experts.
This points us to the so-called embedding-based and model-based measures.
5. What kind of measure do we actually want?
We want thus a distance function
such that: d(A,B)
If two texts express the same idea in different words → small distance.
If two texts differ theologically, logically, or tonally → larger distance.
If two texts are literally identical → distance = 0.
And we want it to be:
Objective (computable, not “feeling-based”)
Sensitive enough for experts
Still reflecting lay-reader equivalence
That suggests something like this:
(A) Semantic embeddings (meaning space)
Instead of comparing word sets, we compare meaning vectors.
Take a good sentence embedding model (which maps text → vector that encodes meaning).
Compute cosine similarity between the two vectors:
This has the property we want:
Free paraphrases with same meaning → vectors very close → small distance
Texts with different meaning or emphasis → vectors diverge → larger distance
This already does what naive readers do:
it says “these are basically the same thing” even if the words differ a lot.
(B) Stylistic / doctrinal layer (expert sensitivity)
But we also want to reflect what an expert sees:
tone differences
theological precision
Qur’anic rhetorical structure
subtle loss or addition of meaning
For that, we can add a second component: style/doctrine distance.
Concretely:
Extract features like:
rhetorical structure (questions vs statements)
emphasis (which clauses are foregrounded)
sentiment / formality / authority tone
presence / absence of key theological terms (stray, guide, rebellious, etc.)
Or—more powerful—train a specialised model (classifier or embedding) on:
pairs of “faithful paraphrase vs distorted paraphrase” labeled by an expert.
Then the model learns an expert-equivalence metric.
So we’d have:
which is small only when both meaning and doctrinal nuance are preserved.
(C) Composite metric: reader vs expert
Then define:
For a naive reader,
dominates,\alpha small.\beta For a specialist, we care more about β.
This gives:
Two paraphrases that feel equivalent to most readers → low
→ overall smalld sem d_\text{sem} .d Two paraphrases that look similar but change a doctrinal nuance (e.g. who guides whom, or whether Allah “allows” vs “wills”) →
jumps → overalld expert d_\text{expert} bigger.d
Now we have a metric that is both objective and expert-sensitive.
(D) Where does “zero distance” fit in?
Under any such metric:
Distance = 0 only if the representations are identical under that metric.
For Jaccard: exactly same word set.
For embeddings: numerically identical embedding (practically, exact same text or trivial variants).
For our composite metric: same meaning + same style + no doctrinal difference.
So:
For Qur’anic Arabic vs its English paraphrases → distance is never exactly zero, but it may be very small under some coarse semantic metric.
For Qur’ān vs LLM-rewritten pseudo-Qur’ān → distance will stay bounded below a certain ε under a good doctrinal/style-aware metric.
That non-zero minimum is our metaphysical “gap”, our “Planck-like cutoff”.
6. Semantic embeddings: Semantic Planck cutoff
Instead of comparing word sets, we map now the whole sentence / paragraph to a vector in semantic space:
The embedding model (a transformer) is trained so that:
Sentences with similar meaning → vectors close together
Sentences with different meaning → vectors far apart
Then we define semantic distance as:
→ semantically very close (almost same meaning).d sem ≈ 0 d_{\text{sem}} \approx 0 moderate → related but different.d sem d_{\text{sem}} large → different meaning.d sem d_{\text{sem}}
This is much closer to “how a reader feels” than Jaccard.
So, with our mosquito verse:
O (original) vs R₁ (tight paraphrase) → expect very small
.d sem d_{\text{sem}} O vs R₃ (freer paraphrase) → still small, but a bit larger.
O vs some random unrelated text → high
.d sem d_{\text{sem}}
Now the metaphysical question becomes:
Can we ever get
for an LLM recreation of O? d sem = 0 d_{\text{sem}} = 0
Or is there always a small ε > 0, even for the best paraphrase?
That’s the semantic version of our “Planck cutoff”.
We can do this very straightforwardly in Python with a sentence-embedding model such as sentence-transformers/all-MiniLM-L6-v2 or similar.
Explicitly, we get
R1: similarity = 0.7986, distance = 0.2014
R2: similarity = 0.8470, distance = 0.1530
R3: similarity = 0.7841, distance = 0.2159
So:
R2 is closest to the original in semantic space
(smallest distance ≈ 0.153)Then R1 (distance ≈ 0.201)
Then R3 (distance ≈ 0.216)
All three are quite close (sim ~0.78–0.85), but none reach 1.0, so none have distance 0.
Thus, even with a faithful paraphrase, the semantic embedding distance is non-zero. In other words, for this verse and this embedding model, LLM paraphrases live at a non-zero distance ε from the source.
This is actually very instructive.
R1 was very literal: “does not hesitate… example… by this trial He lets many go astray…”
R2 was a bit freer: “not embarrassed… give a parable… by means of this, He allows many to wander…”
The embedding model is not measuring “literal closeness”; it’s measuring global semantic + rhetorical similarity in its own learned space.
R2 may be slightly closer in its overall phrasing distribution to the kind of religious / explanatory English the model has seen in training. So in the high-dimensional semantic space, it lands nearer to O’s vector.
The important thing is:
All three paraphrases are semantically in the same cluster as the original.
But none of them coincide with the original in embedding space.
This is exactly our metaphysical intuition:
to a non-expert, all three “feel like” the same verse.
To the embedding, they’re close but not identical.
From just this tiny experiment:
Distance is never zero for paraphrases — even in semantic space.
That’s our semantic ε, a Planck-like cutoff in meaning-space.
Jaccard distance had a lexical ε→ hard lexical cutoff (word-level Planck length):
any word substitution → distance > 0.
It guarantees: no mimicry is literally identical at the word-set level.
Embedding distance gives a meaning-level ε→ softer semantic cutoff (meaning-level Planck length):
free paraphrases can get quite close (0.15–0.20),
but still don’t collapse to 0.
Mimicry can be close, but not identical, even in “meaning space.”
12. Expert metric
1. “Proposition preservation” metric (expert idea → fixed formula)
We formalize “expert concerns” as atomic propositions, then treat everything else mechanically.
For the mosquito verse, we can encode its core content as 3–4 propositions, e.g.:
P1: Allah is not reluctant/shy to give a parable of something as small as a mosquito or smaller.
P2: Believers recognize this parable as truth from their Lord; disbelievers question what Allah means by it.
P3: By means of this parable/test, Allah leads many astray and guides many,
P4: No one is left astray except the rebellious.
Now, for any candidate paraphrase
Then define the expert distance:
If all propositions are preserved → sum = N →
.d expert = 0 d_{\text{expert}} = 0 If half are preserved →
, etc.d expert = 0.5 d_{\text{expert}} = 0.5
Where is “our judgment”?
Only once: in how you choose
.P_1,\dots,P_N After that, everything is mechanical.
In practice, checking “entails/contradicts” can be done with an NLI model (natural language inference): it takes (R, P_i) and outputs a probability of entailment vs contradiction.
Then you define:
with fixed thresholds
2. Structural / role-preservation metric (who does what to whom)
Even more “physics-like”: focus only on roles and relations, not propositions written in English.
For the same verse, we extract from the original:
Entities:
.A = Allah , B = believers , C = disbelievers , D = rebellious A = \text{Allah}, B = \text{believers}, C = \text{disbelievers}, D = \text{rebellious} Predicates / relations in some canonical form, for example:
use_parable(A, mosquito_or_smaller)affirm_truth(B, parable)question_intent(C, parable)cause_stray(A, many)cause_guide(A, many)restrict_stray(A, D)(no one strays except D)
Now for a candidate paraphrase
Parse it (dependency / SRL / information extraction).
Extract its own set of predicate-argument triples
.S R S_R Compare with original’s set
.S O S_O
Define:
Then:
If all roles/relations are preserved →
.d struct = 0 d_{\text{struct}} = 0 If some are missing or altered → distance grows.
Again: the only “expert” step is deciding which predicates/roles you care about.
Once fixed, it’s just: parse → extract → match → compute.
3. NLI-based “symmetry” metric (no hand-crafted features at all)
If we want no explicit propositions, there is a more brutal approach using only NLI:
Treat the original verse
Compute an entailment score
: how muchE ( O → R ) entails .Compute
: how muchE ( R → O ) entails .
Use a pre-trained NLI model; you don’t specify features manually.
Then define:
If they are semantically equivalent in the NLI sense → both entail each other strongly → distance ≈ 0.
If there’s loss, addition, or contradiction → one direction weakens → distance grows.
This is the closest thing to “expert in a box”: the NLI model encodes a lot of world and linguistic knowledge, but you are not manually scoring anything. You just call it.
4. NLI-based, fully automatic “expert”
Recall the idea:
For original text
Use a Natural Language Inference (NLI) model to estimate:
= probability that O entails RE ( O → R ) E(O \to R) = probability that R entails OE ( R → O ) E(R \to O)
Define a symmetric similarity:
Then define a distance:
If they fully say “the same thing” in the NLI sense → both entail each other strongly → sim ≈ 1 → distance ≈ 0.
If some meaning is lost/added/contradicted → at least one direction is weak → sim smaller → distance larger.
This is:
fully automatic (no hand marking),
used elsewhere (NLI + semantic textual similarity / paraphrase tasks),
and can be plugged into your Imru’ al-Qays vs Fātiḥa experiment.
People use NLI models like
roberta-large-mnli,deberta-large-mnli, etc., trained on MultiNLI/SNLI-style datasets, to judge entailment / contradiction.They also use related metrics like BERTScore, BLEURT, and COMET which are all essentially “learned semantic similarity” models built on transformers.
For NLI, we’ll want
transformers(HuggingFace) and maybetorch(it will install as a dependency).These models take
(premise, hypothesis)and give probabilities for: entailment, neutral, contradiction.
Here, what we get:
So:
R2 is best:
→sim ≈ 0.9938 \text{sim} \approx 0.9938 d NLI ≈ 0.0062 d_{\text{NLI}} \approx 0.0062
Then R1:
d NLI ≈ 0.0149 d_{\text{NLI}} \approx 0.0149
Then R3:
d NLI ≈ 0.0208 d_{\text{NLI}} \approx 0.0208
Interpretation:
The NLI model thinks all three paraphrases are essentially logically equivalent to the original.
Differences are there, but they’re at the 0.6–2% level in this metric.
That’s exactly what we wanted from an “AI expert”:
it treats them as near-perfect paraphrases of the same proposition.
Compare with the embedding distances you computed earlier:
d sem ( R 1 ) ≈ 0.20 d_{\text{sem}}(R1) \approx 0.20 d sem ( R 2 ) ≈ 0.15 d_{\text{sem}}(R2) \approx 0.15 d sem ( R 3 ) ≈ 0.22 d_{\text{sem}}(R3) \approx 0.22
So:
Embeddings: “they’re close, but differences still ~15–20% in this space.”
NLI: “they’re almost exact paraphrases (≤ 2% difference).”
We now see the hierarchy:
Lexical (Jaccard) → very harsh, any word change → big difference.
Semantic embeddings → smoother, but still sensitive to stylistic shifts.
NLI → collapses everything that is truly propositionally equivalent into the same cluster; treats R1–R3 as “basically the same statement”.
13. Towards a “Balāgha-Aware” Metric
We’ve measured
lexical similarity (surface words),
semantic similarity (embeddings),
logical / factual similarity (NLI),
and all of that is still “below” what we care about: balāgha — higher-order eloquence, naẓm, rhetorical force.
However, there is no off-the-shelf “balāgha-transformer” that says: “this paraphrase is weaker in rhetoric / iʿjāz” in a principled way.
But we can outline what could be done, and what proxies exist.
1. Why NLI & embeddings are not balāgha
We already see it:
NLI: collapses everything that preserves truth-conditions.
It treats different registers, tones, and rhetorical power as identical if the propositions match.
Embeddings: capture overall meaning + style to some extent, but:
they’re trained to predict context, not evaluate aesthetic quality.
Neither knows:
Is this phrase more concise (ījāz) or bloated?
Is the word order creating surprise, emphasis, rhythm?
Is the sound pattern (sajʿ, assonance, alliteration) stronger here?
Is there a subtle semantic tension or multi-layered metaphor?
That’s what Arabic balāgha studies:
ḥaqīqa–majāz, kināya, taqdīm–taʾkhīr, ījāz–iṭnāb, jinās, sajʿ …
All beyond simple “same meaning / different meaning.”
So of the three, semantic distance is best for “similitude,” but balāgha sits above even that.
2. Do we have a “balāgha transformer” today?
Short answer: no, not in the strong sense we want.
There are things related to it:
Style classification models:
“This text is Shakespearean / Quranic / modern news / poetry / hadith-like.”
Quality/evaluation models:
BLEURT, COMET, GPT-based judges for “goodness” of a translation or summary.
Arabic NLP resources:
Pretrained Arabic BERT/transformers (AraBERT, CAMeL-BERT, etc.),
Some work on Quranic Arabic parsing and rhetorical devices,
But not a fully trained “eloquence score” model.
Nothing like:
Give me two Arabic verses and I’ll output:
0.95 balāgha similarity, 0.3 structural naẓm distance,
and this one is linguistically superior.
That does not exist as a standard tool.
3. What could a “balāgha-aware” metric look like (in principle)?
If we were to design one (this is the interesting part!), we’d need features that correlate with eloquence as understood by classical balāgha. For Arabic, this might include:
Phonological / prosodic features
Patterns of:
consonant clusters,
vowel harmony,
sajʿ (rhymed prose),
recurring end sounds / internal rhyme.
You can extract:
last syllable patterns,
consonant–vowel skeletons (CV patterns),
measures of repetition and symmetry.
Syntactic and word-order features
Frequency and pattern of:
fronting (taqdīm) for emphasis,
ellipsis,
repetition,
parallel structures.
These can be approximated with dependency parsing and counting patterns like:
how often the object comes before the verb, etc.
Figurative density
Harder to automate, but we can approximate:
density of metaphor-related lexical fields,
unusual collocations (words that rarely appear together, signaling creative usage),
shifts from literal to non-literal contexts.
Compactness vs. redundancy
Balāgha is very sensitive to ījāz vs. iṭnāb:
how much meaning is packed per token.
We can approximate “semantic content per word”:
ratio of content words to function words,
mutual information between adjacent words,
how much is repeated vs. how much is new information.
Then, you could:
Extract such features for Quranic verses,
Extract the same for imitations,
Train a model (or just do unsupervised clustering) to see:
whether Quran occupies a distinct region in this “balāgha feature space”,
whether paraphrases always sit at a measurable distance.
This is very much research-level work, not plug-and-play.
4. A practical compromise you could use now
Given what exists today, you can still do something meaningful in the spirit of balāgha using existing LMs:
Language model perplexity
Train or finetune a strong Arabic language model on classical high-register Arabic, including:
Qur’an (if you dare),
early poetry,
classical prose.
Then:
feed it Quranic verses and candidate imitations,
compare per-token perplexity (surprisal).
The hypothesis:
Quranic verses sit near an “optimal” balance between predictability and surprise.
Many imitations will either be too banal (too predictable) or too weird (too high perplexity).
This won’t measure beauty directly, but:
it can detect that Quranic style is statistically balanced in a very special way.
Style classifier
Train a classifier on:
Quran vs. classical poetry vs. hadith vs. modern prose.
Ask it:
“How Quran-like is this sentence?” → probability .
For iʿjāz:
see if genuine Quranic verses form a tight cluster with very high probability,
and whether any generated imitation ever reaches the same level.
This is still far from true balāgha, but it’s at least:
automatic,
Arabic-specific,
and closer to rhetorical “feel” than pure NLI.
5. Where our three metrics stand, conceptually
Let’s rank them with the following analogy:
Logic / NLI
⇒ “Does this say the same proposition?”Good for avoiding hallucination.
Blind to style and eloquence.
Semantics / embeddings
⇒ “Is this roughly the same meaning in a similar context?”Picks up some style.
Still mostly about content, not aesthetic force.
Balāgha / eloquence metric (future work)
⇒ “Is this of similar rhetorical power, compactness, sound, naẓm?”This is the one you really care about for iʿjāz.
Currently no standard transformer-based metric, but possible to design.
In summary, if someone is trying to reproduce Quran, the logical structure is in principle the easiest part to match,
semantic proximity is harder,
and balāgha—if iʿjāz is real in this sense—is where imitation should consistently fail.
There are:
NLI models for logical equivalence,
embedding models for semantics,
style/quality models that approximate some aspects of eloquence.
To actually model balāgha, we’d need to:
define and extract higher-order linguistic features tied to classical rhetoric,
train or at least analyze with them,
then see if Quran vs imitation shows a structural gap in that space.
Quran/iʿjāz and consciousness as problems about higher-order structure in representation space, beyond meaning and logic.
14. Training a Style Model
1. Style Classifier
We also want a model that, given a text (original or regeneration), says something like:
“90% Quran (translated), 5% generic religious English, 5% generic prose”
or “80% Shakespearean, 15% modern fiction, 5% news”
or in Arabic: “Quranic / hadith / classical poetry / modern prose / etc.”
That’s a style classifier.
So:
Input: a text segment
Output: one of a small set of labels (Quranic, Hadith-like, Modern News, Shakespeare, etc.)
The model learns which lexical, syntactic, and higher patterns correlate with each label.
So at test time, you feed:
the original text and
the regenerations
and look at:
orp ( Quranic ∣ text ) p(\text{Quranic} \mid \text{text}) p ( Shakespeare ∣ text ) p(\text{Shakespeare} \mid \text{text})
We then compare how close the regenerations get to the original’s style.
2. Steps to train any style model
Same pipeline for all:
Step 0 – Choose labels
Examples:
For English experiment:
["shakespeare", "modern_fiction", "news"]
For Islamic-Arabic experiment:
["quran", "hadith", "classical_poetry", "modern_arabic_prose"]
For our specific first target:
["quran_translation", "generic_english_religious_prose"]or binary:
["quran_translation", "non_quran"]
Step 1 – Build a labelled dataset
We need lots of short segments with known style:
Quran (Arabic):
each āyah as one sample (or half-ayah if long).
Hadith (Arabic):
each hadith as one sample.
Classical poetry:
single bayt or 2–3 lines.
Modern prose / news:
sentences / short paragraphs from newspapers, blogs, novels.
For English “Quran translated vs other religious/English”:
Class 1: verses from one consistent Quran translation (e.g. Pickthall, Sahih, etc.).
Class 2: generic religious prose (sermons, Christian Bible translations, commentary), plus generic English prose.
We then put them in a table (CSV).
Aim: at least a few thousand examples per class to get something decent.
Step 2 – Pick a base model
You have two main options:
English-only:
e.g.
"roberta-base","bert-base-uncased".
Multilingual (Arabic + English in the same model):
e.g.
"xlm-roberta-base".
For pure Arabic style classification, people often use AraBERT / AraELECTRA–like models (just as examples).
Step 3 – Fine-tune as a classifier
That’s it: we now have a model that takes text and outputs a style label.
3. Zero-shot “Is this Qur’anic style or not?” using the NLI model we ALREADY downloaded
We already pulled microsoft/deberta-large-mnli.
That model can be used as a zero-shot classifier:
Give it a text + some candidate labels expressed in natural language,
and it tells you which label is most entailed.
So we can do:
Premise = your text (original or regeneration).
Hypotheses like:
“This text is an English translation of a verse from the Qur’an.”
“This text is generic English religious prose.”
(or Shakespeare / news / whatever you want)
Then we use entailment probabilities as “style scores”.
How it works conceptually
For each label
: e.g. “This text is an English translation of a verse from the Qur’an.”H_L
We feed (premise = TEXT, hypothesis = H_L) to the NLI model and read P(entailment).
High
P(entailment)⇒ model believes “TEXT is Qur’anic translation” is true.Low ⇒ not so much.
Then we normalize over all labels.
No training. Just NLI.
Running a minimal Python skeleton, reusing our DeBERTa MNLI, will give us style probabilities for each version.
Then our “style distance” can be something like:
So we can directly compare:
original O vs regenerations R1, R2, R3,
Imru’ al-Qays originals vs their imitations,
etc.
This is fully automated and uses a model we already have.
We get:
O {'quran_translation': 0.8524, 'religious_prose': 0.0406, 'news': 0.1069}
R1 {'quran_translation': 0.8818, 'religious_prose': 0.0483, 'news': 0.0699}
R2 {'quran_translation': 0.9485, 'religious_prose': 0.0049, 'news': 0.0467}
R3 {'quran_translation': 0.9196, 'religious_prose': 0.0243, 'news': 0.0561}
So the zero-shot NLI “style classifier” is saying:
All four texts are strongly “Qur’an-translation-like” (0.85–0.95).
The paraphrases R1, R2, R3 are, if anything, even more “Qur’an-translation-like” than the original O in this metric.
Especially R2: ~0.95.
Because what it cannot do (in this form):
Recognize “this is the exact original translation vs a paraphrase”.
Rank “this is the true Qur’an vs an extremely good imitation” — because we didn’t give it that kind of training signal.
Right now, it’s more of a genre detector than a balāgha/iʿjāz detector.
Indeed, what we were asking is a generic question:
“Is this text Qur’an-translation-like in general?”
But what we want really to ask is conditional:
“Given that O is Qur’an, how likely is it that R1, R2, R3 are also Qur’an (same source / same corpus) rather than just Qur’an-style paraphrases?”
These are really two different problems.
4. Unsupervised “closeness to corpus” using sentence embeddings (using sentence-transformers we ALREADY installed)
This is actually simpler:
Pick a corpus for each style:
A bunch of Qur’an translation verses (class Q).
A bunch of other religious prose (class R).
(Similarly: Shakespeare vs modern fiction vs news.)
Use
SentenceTransformerto encode all texts into embeddings.Compute the centroid (average embedding) for each class:
𝑐 𝑄 𝑐 𝑅
For any new text
:Embed it →
.𝑒 𝑇 Compute cosine similarity with both centroids:
cos ( 𝑒 𝑇 , 𝑐 𝑄 ) cos ( 𝑒 𝑇 , 𝑐 𝑅 )
Define a “Qur’anic style score”:
This gives you a continuous score between 0 and 1: “how much this text lives in the Qur’an-translation region of embedding space”.
Again: no training, just classical vector arithmetic.
Neither is true balāgha, but:
Zero-shot NLI classifier:
more discrete (“does this sound like Qur’an or like news?”),
uses the logical/semantic NLI brain as a proxy for style + content.
Embedding-centroid method:
more continuous,
may capture some stylistic nuance simply because the embedding model has seen lots of data.
We can even combine them:
Use embeddings for a semantic style axis,
Use NLI zero-shot scores as a label-style axis.
No dataset labeling.
No fine-tuning.
You rely on pretrained models (which we already have) and simple math:
cosines,
softmax-normalized entailment scores.
Turn DeBERTa-MNLI into a zero-shot style classifier.
Use sentence-transformers to do corpus-style similarity without any extra training.