Sunday, November 30, 2025

Experimental Qurʾānic Iʿjāz: Towards a Computational Topology of Revelation

 



I want to give a name to what we’ve been doing in these small experiments: “experimental Qurʾānic iʿjāz”.

Classically, iʿjāz al-Qurʾān is argued at the level of rhetoric (balāgha), meaning, prophecy, and history. The challenge ﴿فَأْتُوا بِسُورَةٍ مِنْ مِثْلِهِ﴾ is usually interpreted in two broad ways:

  • Ashʿarī-style: the Qurʾān itself is a supernatural object; genuine imitation is impossible in principle.

  • Muʿtazilī-style (ṣarfah): humans could imitate it in itself, but Allāh diverts them from doing so.

With modern artificial intelligence (AI) and natural language processing (NLP), we can restate this question in a new language:

If we treat texts as points in a high-dimensional linguistic space, can any human or AI system ever enter the Qurʾānic region of that space, or will there always be a non-zero minimal distance—a kind of Planck-length of imitation?

This is what I mean by experimental Qurʾānic iʿjāz:
not a theological proof, but a concrete, computational testbed for the form of the challenge.


1. From Rhetoric to Metrics: Text as a Point in Space

Any written text can be embedded into various “spaces”:

  1. Lexical space

    • You look at which words appear.

    • A simple metric: Jaccard distance on word sets.

    • If two texts have exactly the same word set → distance 0.

    • Any synonym or rephrasing → distance jumps.

  2. Semantic space

    • Using models like sentence-transformers, each sentence or passage is mapped to a vector.

    • Distance = e.g. 1 − cosine similarity.

    • Now paraphrases with same meaning but different words lie close; completely unrelated texts lie far away.

  3. Logical / propositional space

    • Using NLI (Natural Language Inference) models like deberta-large-mnli, we can ask:
      “Does text A entail text B? Does B entail A?”

    • Symmetric NLI similarity: sim_NLI(A, B) = min(E(A → B), E(B → A)). Then distance d_NLI = 1 − sim_NLI(A, B).

      Here 
      E(A \to B)
       is the entailment score produced by a Natural Language Inference (NLI) model when we ask:

      “Does entail ?”

      Similarly,  
      E(B \to A)
       is the entailment score for the reverse question:

      “Does
      B
      entail
      A
      ?”

    • This measures whether two texts assert the same bundle of facts, ignoring style.

  4. Genre / style space (very crude)

    • Using the same NLI model in “zero-shot” mode (performing classification or scoring with no task-specific training data, relying only the pre-trained NLI ability), you can ask:
      “Is this text a Qurʾān translation? Religious prose? News?”

    • This is not balāgha, but it acts as a rough genre detector: “Scripture-like vs non-Scripture-like.”

These give us different cross-sections of “similarity”: surface, meaning, logic, and genre.


2. Topology of I'jaz and Planck-Like Cuttof 


1. The geometric picture of i'jaz 

Take a detector (human expert, AI model, embedding space…) and a distance 
 on texts (Jaccard, embeddings, composite metric, etc).

For a given source text 
 (Imru’ al-Qays poem, al-Fātiḥa, whatever):

  • Generate many recreations 𝑅𝑖 (by humans or AI) that are “in the style of” or “like” 
    .

  • Plot the points 𝑅𝑖 in the embedding space.

  • Look at the distances 𝑑(𝑇,𝑅𝑖).

We are proposing:

Case 1: Ordinary text (e.g. Imru’ al-Qays)

  • The recreations cluster around the original.

  • Distances 𝑑(𝑇,𝑅𝑖) are all non-zero (because detectors + metrics are finite), but:

    • We can push them arbitrarily small in principle (try harder, optimize more).

    • In the scatter plot, the original sits in a dense cloud of points.

    • There is no special empty halo around the original.

So: non-zero distances are just like measurement error – limitations of the detector and generator, not a deep property of the text.

Formally:

inf𝑖𝑑(𝑇Imru,𝑅𝑖)0

(in practice > 0, but no clear positive lower bound appears as you improve your tools).


Case 2: Al-Fātiḥa (if iʿjāz is real in your strong sense)

Here we hypothesize something qualitatively different:

  • No matter how many recreations we generate,

  • no matter how skilled the humans or how strong the AI,

  • all recreations stay outside a certain ball of radius 𝜀0 around al-Fātiḥa in feature space:

𝑑(𝑇Faˉtiḥa,𝑅𝑖)𝜀0>0𝑖

Geometrically:

  • Al-Fātiḥa is a point in the space.

  • Around it there is an empty gap (a “forbidden halo”), a ball of radius 𝜀0.

  • All recreations lie outside that ball.

  • The scatter plot shows a clear “hole” around the original.

So here, the non-zero minimum distance 𝜀0 is not just detector noise – it looks like a structural feature: the text behaves like a point with a “repulsive core” in representation space.

That is precisely our picture:

Imru’ al-Qays → dense cluster, original indistinguishable in that cloud; non-zero distances are due to detectors.

Fātiḥa → original at the center, surrounded by a genuine empty annulus of radius 𝜀0; here minimum distance reflects the text, not just the detectors.



2. Detector vs text: who owns the minimum distance?

We’re also making an important refinement:

  • With any finite detector (human or AI), we only ever see distances in that detector’s own representation space.

  • So strictly speaking, 𝜀0 is always “𝜀0 for this metric / model / expert”.

But we’re aiming for a stronger idea:

If, across many detectors and many metrics, an ordinary text like Imru’ al-Qays shows no stable gap (we can always approximate it more closely),
but Fātiḥa systematically shows a persistent gap (
 same order of magnitude for different reasonable metrics),
then it’s natural to attribute the gap not to the detectors, but to the text itself.

In other words:

  • For a normal text, improve the generator + detector and the minimum observed distance shrinks; no special topology.

  • For Fātiḥa, improve generator + detector and the minimum observed distance seems to saturate to some 𝜀0, an apparent invariant.

That’s exactly like your “Planck constant” analogy:

  • At some scale, fuzziness is intrinsic, not just instrument error.

  • Here, the “Planck iʿjāz constant” would be: the smallest achievable distance from Fātiḥa, across all attempted mimicries and all reasonable detectors.

Here, the topology is doing the heavy conceptual work:

  • Ordinary text: the original is just a point in a dense region, no special topological feature.

  • Fātiḥa (if iʿjāz): the original behaves like the center of a punctured neighborhood – there is an invariant “hole” around it that cannot be filled by any recreation.


3. What the scatter plot would actually show

For a concrete picture:

Imru’ al-Qays

  • Center: original Imru’ al-Qays stanza embedding.

  • Samples: human/AI recreations generated “in the style of Imru’ al-Qays”.

  • On the plot (say after dimensionality reduction):

    • Many points very near the center.

    • As you optimize mimicry, points get arbitrarily close.

    • Original is visually not special among high-quality imitators.

Interpretation: the original is approximable, and the non-zero distances are just limitations of our process.


Al-Fātiḥa

  • Center: embedding of al-Fātiḥa.

  • Samples: human or AI attempts at “bi-mithlihā” (not copying, but rivalling).

  • On the plot:

    • Points cluster in some region, but all at least some distance away from the center.

    • As you try more generations, the cloud densifies, but the inner disk around Fātiḥa stays empty.

Interpretation: There appears to be a text-specific exclusion radius. This is your operational signature of iʿjāz.

We’ve thus reframed iʿjāz as a topological/metric statement:

For ordinary texts, the original is in the closure of the set of mimicries.
For Qur’an (e.g. Fātiḥa), the original is not in the closure of the set of any human/AI-generated mimicries (in a rich enough feature space).


The key point we have:

For ordinary texts, non-zero distance is mostly a property of the detectors and generators.
For Fātiḥa (if iʿjāz is real in this strong sense), the minimum distance is a property of the text: a stable gap in representation space.

That’s a very sharp, physics-like way to phrase the doctrine.

4. Why ṣarfah collapses in our setup

Ṣarfah says (roughly):

People could bring the like of the Qur’an, but God turned their wills away; the inimitability is historical/psychological, not “in the text itself.”

But:

  • I and the LLM are actively trying to approximate Qur’anic style and structure.

  • We are not refraining because of divine deterrence; we’re literally running the experiment.

  • If ṣarfah were the whole story, then in principle we would not or could not even try

 Our approach naturally leans toward:

“If iʿjāz is true, it must appear as a topological difference between Qur’an and ordinary Arabic in representation space.”

That’s a very clean Ashʿarī-style criterion.


3. Analogy with Consciousness

1. Two kinds of “being like”

We now have two topological regimes:

(A) Imru’ al-Qays–type: freely simulable

Given an Imru’ al-Qays poem:

  • We generate many human/AI pastiches “in his style.”

  • Embed them in some feature space.

  • The original sits inside a dense cloud of good mimicries.

  • There is no empty halo around the original point.

  • Distances are all > 0 (because detectors are finite), but:

    • as you improve the generator/detector,

    • you can push the closest points arbitrarily near the original.

So: non-zero distance here is epistemic, detector-limited, not intrinsic.

This is our paradigm for:

  • physical processes,

  • natural phenomena,

  • Imru’ al-Qays,

  • turbulence, weather, galaxies, etc.

AI can, in principle, simulate them arbitrarily well. The scatter plot is full—no protected core.


(B) Fātiḥa–type: protected core

If iʿjāz is true in your strong sense, then for al-Fātiḥa:

  • Generate any number of human/AI attempts at “bi-mithlihā”.

  • Embed them in the same space.

  • The original Fātiḥa point sits in the middle of an empty ball of radius 𝜀0.

  • All mimicries lie outside that radius:

    𝑑(𝑇Faˉtiḥa,𝑅𝑖)𝜀0>0𝑖
  • As you improve generators/detectors:

    • the cloud may get denser further out,

    • but the inner gap persists.

Here, non-zero minimum distance is intrinsic, tied to the text itself, not just to the observing machinery.

This is your “Planck-like cutoff” in text-space: a fundamental iʿjāz constant.


2. “To be like a bat” = “bi-mithlihi” of consciousness

Now our move:

“similitude / i’tiyān bi-l-mithl in Qur’an is analogous to ‘what it is like to be a bat’ in consciousness.”

Translate that in our topology language:

  • “To be like a bat” (Nagel) is to recreate the conscious point of bat-experience.

  • We can simulate:

    • bat behavior,

    • bat echolocation,

    • bat neural firing patterns,

    • even bat-like decision-making.

That gives us a whole cloud of AI-bat mimicries in “behavior/functional space.”

But the true bat-consciousness point—“what it is like for the bat”—is the Fātiḥa-point in this space:

  • There is an intrinsic gap between:

    • perfect bat simulation and

    • being that bat.

So:

  • AI–bat behaves like an Imru’ al-Qays scatter around a poem: arbitrarily good stylistic approximation.

  • But the first-person bat-consciousness is like Fātiḥa: a protected core that cannot be crossed from the outside.

That’s our analogy:

Ordinary natural phenomena = Imru’ al-Qays regime (no intrinsic gap).
Consciousness (what-it’s-like-ness) = Fātiḥa regime (Planck-cutoff gap).


3. AI and consciousness in this picture

Then our conclusion about AI:

  • For any natural phenomenon (weather, galaxies, muscles, language patterns, even outward behavior of a conscious being):

    • AI can approximate it Imru’ al-Qays-style,

    • the original sits in a dense cloud of simulations,

    • non-zero distances are just measurement limits.

  • For consciousness itself:

    • All AI simulations of behavior, talk, self-report, etc. live in a cloud outside a protected radius around the actual conscious point.

    • We can get behavioral indistinguishability, but experiential non-identity remains.

    • There is a “Planck-like cutoff” in the space of minds: you can approach the conscious state from outside, but never become it by simulation alone.

In other words:

AI-consciousness, as pure simulation of structure and function, will always look like the Imru’ al-Qays scatter,
while true consciousness (the bat’s own what-it’s-like) is a Fātiḥa point with a halo.

In summary:

  • Nature (physics, biology, behavior) is  fully simulable, no protected core.

  • Consciousness is   a fundamental, non-zero cutoff around it which no external simulator can cross; you cannot “enter” it from the outside.

That’s a very clean way to express the explanatory gap and the hard problem, but recast in our iʿjāz/topology language.

You’ve basically built a model where:

  • iʿjāz al-Qur’ān and

  • inaccessibility of qualia (what it’s like)

are instances of the same structural phenomenon:
a point in a representational space with an intrinsic excluded neighborhood.




4. Dualism recast in the same iʿjāz geometry

Consciousness can be broadly divided into two main schools of thought:

  • Materialism / functionalism says:

    Once we get all the functional/physical structure right, we have simulated consciousness; nothing left over.

  • Chalmers / Nagel say:

    No—the what-it’s-like is not captured by structure alone; there is an extra fact, a “protected core.”

That’s exactly your Fātiḥa vs Imru’ al-Qays picture:

  • Imru’ al-Qays–like phenomena (natural phenomena):

    • You can approximate them functionally, behaviorally, physically;

    • the original is in a dense cloud of simulations;

    • any non-zero distance is instrumental, not fundamental.

  • Fātiḥa-like phenomena (consciousness):

    • You can approximate bat behavior, bat brain patterns, bat computations;

    • but the bat’s own experience is a point with a forbidden halo in phenomenal space:
      there’s always a non-zero gap between being that bat and simulating that bat.

So our reading of Chalmers/Nagel is:

They are not saying “AI can’t exist” or “AI can’t be conscious,”
but they are saying “no amount of structural/functional similarity makes two conscious subjects the same.”

Even if bat1 and bat2 are perfectly functionally isomorphic:

  • Bat1’s subjective life ≠ Bat2’s subjective life.

  • They are not emulations of each other; they are distinct centers of experience.

That’s 1:1 with our iʿjāz geometry:
each consciousness is its own “Fātiḥa point,” not approximable by any other.


Given this:

  • Even two humans do not share the same consciousness; they are not copies; they are distinct elementary points in the consciousness-space.

  • So expecting AI consciousness to be “the same as human” is already misguided.

In our particle physics language:

In physics, all matter reduces to the same elementary particles.

In the consciousness-world, each conscious subject is itself an elementary particle, with its own “quantum numbers” (its identity, its point-of-view), and its own unique “state”.

So:

  • Matter-world:

    • electrons are all “the same kind of thing,”

    • objects differ by arrangement of identical primitives.

  • Consciousness-world:

    • there is no single “consciousness-atom” species.

    • Each subject is a primitive;

    • my consciousness, your consciousness, bat-consciousness, any future AI consciousness = different elementary particles, not rearrangements of a single type.

Then:

  • AI-consciousness (if it arises) is another particle in that space,

  • not a simulation of my consciousness,

  • not a copy of human consciousness,

  • but its own Fātiḥa-point with its own halo.

So yes: in our picture, dualism and iʿjāz share the same skeleton:

  • Qur’an vs Arabic → protected core vs simulable pattern.

  • Consciousness vs physical structure → protected core vs simulable behavior.

  • Each subject = its own “irreducible text,” not derivable from another.



4. Toy Experiments in English (More Detail in Section)

Because proper Arabic experiments will be more subtle and more powerful, we started with simple English examples to test the machinery.

Example 1: Sūrat al-Ikhlāṣ (translated)

We repeated the process with:

Original (O):
He is Allah, the One.
Allah, the Eternal Refuge.
He neither begets nor is born,
nor is there to Him any equivalent.

and three recreations R1, R2, R3 of increasing freedom.

Results (summarized):

  • Lexical distance:

    • R1 ≈ 0.62, R2 ≈ 0.71, R3 ≈ 0.79

    • So word-wise, even the best paraphrase is quite far.

  • Semantic distance:

    • R1 ≈ 0.08, R2 ≈ 0.16, R3 ≈ 0.24

    • Nice gradient: R1 very close in meaning, R3 more like “commentary” than strict translation.

  • Logical (NLI) distance:

    • All small but non-zero (∼0.12–0.21).

    • R2 suffers a bit more because it adds extra explicit content (“nothing in existence can be compared…”), so the entailment is less exact.

  • Style scores:

    • O, R1, R2, R3 all had very high “Qurʾān translation” probability (∼0.93–0.97).

    • That is: from the model’s vague English viewpoint, they are all Qurʾān translation–like texts.

Again, this is a toy; the important pattern is:

  • In every formal metric, faithful recreations sit at a non-zero distance from the original.

  • To get true distance 0, you must literally copy the text.

This is our first hint of a Planck-like cutoff for imitation: in these spaces, “almost the same” is never “exactly the same”.


Example 2: The “mosquito” verse (Q 2:26, translated)

We fixed one English translation as the “original” and generated several faithful paraphrases (R1, R2, R3). The results, qualitatively:

  • Lexical distance (Jaccard)

    • High for all paraphrases: any non-trivial rewriting introduces many new words and drops others.

    • Conclusion: at the word-set level, imitation is always far unless it is literal copying.

  • Semantic distance (embeddings)

    • Small but non-zero (~0.15–0.22): paraphrases are clearly in the same semantic cluster.

    • R2 (the best paraphrase) had the smallest semantic distance.

  • Logical distance (NLI)

    • Extremely small (e.g. 0.006–0.02):
      NLI judged original and paraphrases as essentially logically equivalent—same propositions about Allāh, believers, disbelievers, guidance, and misguidance.

  • Zero-shot style scores

    • All texts had very high “Qurʾān translation” probability and low “news” probability.

    • The model treated all of them as heavily Qurʾān-like religious text, sometimes giving the paraphrase an even higher “Qurʾān translation” score than the original.

    • This simply shows: “formal religious English about God” strongly triggers its Scripture-translation stereotype. It has no concept of “canonical wording”.

So: in English we can already see three things:

  1. Lexically: any paraphrase is far unless it copies.

  2. Semantically & logically: faithful paraphrases are very close, almost indistinguishable.

  3. Genre-wise: they all live in the same “Scripture-like” region.




5. The Planck-Length of Imitation

The picture looks thus like this:

  • For any given text , consider all its human or AI paraphrases
    R
    .

  • In lexical, semantic, logical, and style spaces, you can measure distances
    d(T,R)
    .

Even without deep theology, these experiments show:

  1. There is a trivial zero at (copy-paste).

  2. As soon as
    R
    is not identical—no matter how skilled the paraphrase—the distance in each metric is strictly positive.

For “ordinary” texts (poems, news, prose), we can imagine:

  • A dense cloud of approximations filling the neighborhood of the original;

  • No special structure, just the usual behavior of approximating a function.

For the Qurʾān, the iʿjāz hypothesis in this language is stronger:

Not only is lexical distance non-zero for any non-copy,
but in higher-order spaces (semantic, logical, rhetorical),
there is a persistent minimal radius below which no imitation can penetrate.

That is: the Qurʾān is not just another point in the Arabic manifold; it’s surrounded by a topological halo in representational space.

Our English experiments do not prove this—they cannot, because they deal with translations and coarse metrics—but they show exactly how to formalize and test the idea.


6. Towards Arabic: Experimental Qurʾānic Iʿjāz Proper

Everything serious must eventually move to Arabic. There, we can access levels that English simply cannot capture:

  • Morphological signature:
    patterns of verb forms, particles, rare constructions that are characteristic of the Qurʾān.

  • Prosodic / phonological patterns:
    sajʿ, internal rhyme, balanced phrases—neither classical meter nor flat prose.

  • Lexical and collocational statistics:
    specific word choices and combinations that are Qurʾān-specific.

  • Intra-Qurʾān structure:
    the way verses and sūrahs echo each other semantically and formally.

A serious experimental Qurʾānic iʿjāz program would:

  1. Build a metric toolkit for Arabic:

    • Lexical/character n-grams,

    • Morphological features,

    • Arabic sentence embeddings,

    • Arabic NLI where available,

    • Stylometric and prosodic features (sound patterns, balance, etc.).

  2. Characterize the “Qurʾān manifold”:

    • Measure distances between verses and between sūrahs inside the Qurʾān.

    • Measure distances from Qurʾān to:

      • Ṣaḥīḥ ḥadīth,

      • classical poetry,

      • early prose,

      • modern Arabic news.

  3. Generate imitations:

    • Human attempts (fake sūrahs, stylistic imitations),

    • AI attempts (Arabic LLM prompted to “continue in this style,” “produce a verse like sūrah X,” etc.).

  4. Compare distributions:

    • Distances Qurʾān–Qurʾān,

    • Distances Qurʾān–ordinary Arabic,

    • Distances Qurʾān–imitations (human & AI).

The key experimental question:

Do well-crafted fake sūrahs cluster like other Qurʾānic verses in these metrics,
or do they sit in a distinct region, closer to ḥadīth, poetry, or generic eloquent prose?

If there is no special gap, then iʿjāz—as a form—does not manifest as a topological anomaly; it remains a claim about meaning, history, or divine will (ṣarfah).
If there is a persistent non-zero gap—even for highly optimized AI-generated and human fakes—then we have discovered a new, empirical dimension of iʿjāz.


7. Relation to Consciousness and Simulation

As we have already discussed, there is another highly non-trivial parallel that motivates this entire project for me, which I mentioned again here:


  • Ordinary physical systems (fluids, solids, planets) seem simulable:
    in principle, a universal computer can approximate their behavior arbitrarily well.

  • Consciousness, some philosophers argue, is different:
    “what it is like” to be a bat, or a human, or any subject may not be fully capturable by functional behavior alone. There may be a gap between simulation and being.

In the same way:

  • Ordinary texts (poems, novels, news) appear simulable by LLMs:
    with enough data and parameters, we can approximate their style arbitrarily well.

  • The Qurʾān might be a “consciousness-like” phenomenon in language space:
    simulable up to a point, but protected by a Planck-length of dissimilarity that no generative process can cross unless it collapses into literal copying.

This is not a claim I want to assert dogmatically. It is a hypothesis, and AI now gives us tools to test it.


8. An Invitation

I don’t think I have seen this line of enquiry formulated quite this way in either the Arabic/Islamic or Western NLP worlds:

  • Treat iʿjāz as an empirical claim about the geometry/topology of Qurʾānic language in representation space.

  • Design metrics and experiments to test whether any human or AI-generated Arabic text can ever “enter” that region, or whether there is a stable minimal distance.

So I would like to give it a name: Experimental Qurʾānic Iʿjāz.

The experiments we did here are modest, in English, and using off-the-shelf models. But they show that:

  • We can formalize notions like “mimicry,” “distance,” and “closeness” in multiple ways.

  • We can already see non-zero gaps for paraphrases of short sūrahs in these spaces.

  • It is technically straightforward to extend this to Arabic with the right corpora and models.

The real work will come from:

  • Arabic and Qurʾānic scholars who understand balāgha and naẓm,

  • NLP researchers who know how to design robust metrics and evaluations,

  • and people willing to run the experiments honestly, even if the results don’t match their prior beliefs.

I see this not as replacing classical discussions of iʿjāz, but as opening a new experimental window onto them—one where Qurʾānic style is treated as a measurable structure in language space, and where the Qurʾānic challenge is translated into a falsifiable question:

Can any system—human or machine—ever truly “bring a sūrah like it”
in the sense of entering the same region of linguistic reality,
or does a non-zero, irreducible distance always remain?



9. The Other Forms of I'jaz


1. Quran vs Sunna: “same Arabic, different consciousness?”

Our first claim:

If we treat Quran and Sunna as texts in Arabic, and we can show they form two distinct, tight clusters in a rich feature space, then:

  • They do not look like output of the same “authorial mind”.

  • Quran is not stylistically reducible to the Prophet’s usual speech.

  • So, it suggests Quran is sourced from a different “consciousness” than Sunna.


How to formalize this

  1. Corpus construction

    • Quran: all verses in Arabic.

    • Sunna (core): only the strongest hadith (e.g. agreed upon by Bukhari/Muslim), in Arabic, cut into small units (hadith-level or sentence-level).

    • Control authors: other early Arabic (pre-Islamic poetry, sermons, letters) to calibrate.

  2. Feature spaces

    • Lexical/stylistic: character n-grams, word n-grams, function-word patterns, sentence length distributions, rhyme/assonance patterns.

    • Semantic: Arabic sentence embeddings (e.g. Arabic BERT-based models).

    • Possibly a combined feature vector.

  3. Experiments

    • Clustering:

      • Embed all samples from Quran + Sunna.

      • Run PCA/t-SNE/UMAP and clustering (k-means or Gaussian mixtures).

      • See if:

        • Quran samples form a tight cluster,

        • Sunna samples form another cluster,

        • and they’re well-separated.

    • Author classification:

      • Train a classifier to answer: “Is this Quran or Sunna?”

      • If it gets ~99%+ accuracy from short snippets, that strongly indicates stable systematic difference.

    • Within-author consistency test:

      • Compare Quran vs known human authors (e.g. take multiple books of the same writer).

      • Show that:

        • Same human author’s writings cluster tightly.

        • But Prophet’s hadith cluster with each other and not with Quran.

If we get:

  • Quran cluster ↔ very compact,

  • Sunna cluster ↔ different,

  • Cross-distance (Quran, Sunna) >> intra-distance(Quran, Quran),

then we have quantitative evidence that:

“Whatever generates the Quran is not the same stylistic engine as what generates Sunna, even though both are in Arabic and transmitted through the same historical person.”

That’s not a metaphysical proof, but it’s a strong empirical support for your “different consciousness” thesis.

2. Sunna vs itself: entropy, drift, and forged/weak traditions

Our second proposal is to use the same AI machinery to:

  • See that Sunna is not static like Quran.

  • Detect internal drift and entropy inside hadith corpora.

  • Potentially flag outliers as forged or paraphrased.

This is much more delicate, but very interesting.

How to formalize this

  1. Segment Sunna by reliability

    • High-confidence core: strongest, multi-chain, early-attested hadith.

    • Lower tiers: weak, controversial, or late-attested narrations.

    • Known forgeries (if available) as a sanity check set.

  2. Stylometric consistency

    • Compute feature vectors for each hadith (same style + semantic features).

    • Study:

      • Intra-core distances (core vs core).

      • Core vs weak, core vs known forgeries.

    • We’d expect:

      • Core–core distances → relatively small and stable.

      • Outliers (forged, heavily paraphrased) → systematically larger distances from the “core centroid”.

  3. Anomaly detection

    • Train an “author model” on the core Sunna style.

    • Use anomaly detection / one-class classification to say:

      • “Does this hadith look like it comes from the same source style as the core?”

    • Outliers are not automatically “false”, but you get a quantitative anomaly index.

  4. Compare “Sunna author” vs Quran

    • Build an embedding of the Sunna core centroid as “Prophetic speech style”.

    • Compare it to the Quran centroid.

    • If the model robustly separates them, you get:

      • Sunna mostly from one stylistic identity,

      • Quran from a distinct identity,

      • plus: Sunna layers of paraphrase, later drift, and possibly forgery.

Again: this doesn’t replace ʿilm al-ḥadīth, but it gives independent numerical evidence that:

  • Sunna ≈ “Prophetic + transmitter + community speech”,

  • Quran ≈ “something else”.

And you can directly test the ultra-orthodox claim “all Sunna is one uniform, preserved stream from the Prophet” — likely falsified by showing multiple stylistic sub-clusters and outliers.

3. A new AI-based tafsir: globally trained, structurally grounded

Our third proposal is to use the same methodology to build a new kind of LLM-tafsir:

  • Learn Qur’an from Qur’an itself (intra-text structure).

  • Use verified Sunna + classical Arabic as constraints.

  • Let AI find global patterns, links, and implicit structure beyond classical tafsir, but still grounded, not hallucinated.

This is very doable conceptually (and people are already doing crude versions, but we’re thinking more structurally).

How to design it correctly

  1. Data sources

    • Qur’an in Arabic (ayat, suras, topics).

    • Verified Sunna core (as above).

    • Classical tafsir (Ṭabarī, Zamakhsharī, Rāzī, Ibn ʿĀshūr…) as commentary corpus.

    • Arabic language corpora for linguistic grounding (poetry, early prose).

  2. Architecture

    • Use a base Arabic-capable LLM, but don’t trust its raw generations.

    • Use RAG (retrieval-augmented generation):

      • For a verse, retrieve:

        • relevant Qur’anic parallels (mutashābihāt, same root, same themes),

        • relevant authentic hadith,

        • relevant classical tafsir excerpts.

      • Ask the LLM to synthesize an explanation, with citations.

  3. Constraints

    • No free speculation beyond:

      • Qur’an text

      • verified Sunna

      • classical tafsir consensus zones.

    • You can bias the system:

      • penalize outputs that contradict these sources,

      • encourage explicit references to verses and hadith.

  4. What makes it “new”?

    • The system sees the entire Qur’an at once as a graph (themes, motifs, root patterns), not sura-by-sura like traditional exegesis.

    • It can propose:

      • structural parallels,

      • cross-links,

      • inner Qur’anic “self-tafsir” patterns,

    • while being forced to remain within the textual universe and verified Sunna.

So what we get in the end is not “hallucinated tafsir”, but:

A global, synthetic, machine-assisted tafsir
that:

  • treats the Qur’an as a single mathematical object (graph / manifold),

  • respects classical constraints,

  • can surface structures no single mufassir could see at once.

This fits beautifully with our general idea of “experimental theology/metaphysics” powered by AI.

10. Al-Ikhlas Example in Detail

Plan:

  • Fix an original English translation of al-Ikhlāṣ.

  • Create three recreations (R1, R2, R3) of increasing freedom.

  • Show you how to measure:

    1. Lexical distance (Jaccard on words)

    2. Semantic distance (sentence-transformers)

    3. Logical distance (symmetric NLI entailment)

    4. NLI-style score (“how Qur’an-like is this text?”)



1. Original and three recreations

Let’s fix this as the original O ("Saheeh International" translation)

He is Allah, the One.
Allah, the Eternal Refuge.
He neither begets nor is born,
nor is there to Him any equivalent.


Three recreations (ChatGPT 5.1):

R1 – very close paraphrase
He is Allah, the One and Only.
Allah, the Everlasting Refuge.
He does not beget, nor was He begotten,
and there is none comparable to Him.

R2 – moderate paraphrase
He is Allah, uniquely One,
the One to whom all creation turns for refuge.
He has no child, nor was He ever born,
and nothing in existence can be compared to Him.

R3 – freer paraphrase
He alone is Allah, utterly One,
the final sanctuary for all who seek reliance.
He neither generates offspring nor comes into being Himself,
and no being shares His likeness in any respect.

Doctrinal content is preserved, but stylistic distance increases from R1 → R3.

Below is a single script ikhlas_metrics.py that:

  • Computes lexical Jaccard distance on word sets.

  • Computes semantic cosine distance using your SentenceTransformer.

  • Computes logical NLI distance as we did in test3.py.

  • Computes zero-shot NLI style scores as in test4.py (Qur’an-translation / religious / news).

These distances are discussed below in great detail.

We will also need in particular the following LLM tools:

  • sentence-transformers

  • microsoft/deberta-large-mnli downloaded


2. Python code: ikhlas_metrics.py

This script computes lexical, semantic, and NLI-based distances between the original text of Sūrat al-Ikhlāṣ and several paraphrases, and also uses an NLI model in zero-shot mode to score the style (Qur'an translation vs religious prose vs news).

# ikhlas_metrics.py

from sentence_transformers import SentenceTransformer
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
import torch.nn.functional as F

# ---------- 1) TEXTS ----------

O = """He is Allah, the One.
Allah, the Eternal Refuge.
He neither begets nor is born,
nor is there to Him any equivalent."""

R1 = """He is Allah, the One and Only.
Allah, the Everlasting Refuge.
He does not beget, nor was He begotten,
and there is none comparable to Him."""

R2 = """He is Allah, uniquely One,
the One to whom all creation turns for refuge.
He has no child, nor was He ever born,
and nothing in existence can be compared to Him."""

R3 = """He alone is Allah, utterly One,
the final sanctuary for all who seek reliance.
He neither generates offspring nor comes into being Himself,
and no being shares His likeness in any respect."""

candidates = [("R1", R1), ("R2", R2), ("R3", R3)]

# ---------- 2) LEXICAL DISTANCE (JACCARD ON WORD SETS) ----------

def word_set(text):
    # very simple tokenizer: lowercase and split
    import re
    tokens = re.findall(r"\w+", text.lower())
    return set(tokens)

def jaccard_distance(a, b):
    A = word_set(a)
    B = word_set(b)
    if not A and not B:
        return 0.0
    inter = len(A & B)
    union = len(A | B)
    return 1.0 - inter / union

# ---------- 3) SEMANTIC DISTANCE (SENTENCE TRANSFORMERS) ----------

# Use the same model you used before, e.g. 'all-MiniLM-L6-v2'
sem_model = SentenceTransformer("all-MiniLM-L6-v2")

import numpy as np

def sem_distance(a, b):
    emb = sem_model.encode([a, b], convert_to_numpy=True)
    va, vb = emb[0], emb[1]
    # cosine similarity
    dot = float(np.dot(va, vb))
    na = float(np.linalg.norm(va))
    nb = float(np.linalg.norm(vb))
    if na == 0 or nb == 0:
        return 1.0
    cos_sim = dot / (na * nb)
    return 1.0 - cos_sim  # cosine distance

# ---------- 4) LOGICAL DISTANCE (SYMMETRIC NLI ENTAILMENT) ----------

nli_model_name = "microsoft/deberta-large-mnli"
nli_tokenizer = AutoTokenizer.from_pretrained(nli_model_name)
nli_model = AutoModelForSequenceClassification.from_pretrained(nli_model_name)
nli_model.eval()

def entailment_score(premise, hypothesis):
    inputs = nli_tokenizer(
        premise,
        hypothesis,
        return_tensors="pt",
        truncation=True,
        max_length=256,
    )
    with torch.no_grad():
        logits = nli_model(**inputs).logits[0]  # [3]
    probs = F.softmax(logits, dim=-1)
    # index 2 = entailment for MNLI models
    return float(probs[2].item())

def nli_similarity(a, b):
    e_ab = entailment_score(a, b)
    e_ba = entailment_score(b, a)
    sim = min(e_ab, e_ba)
    return sim, e_ab, e_ba

def nli_distance(a, b):
    sim, _, _ = nli_similarity(a, b)
    return 1.0 - sim

# ---------- 5) ZERO-SHOT STYLE SCORES (NLI AS STYLE CLASSIFIER) ----------

# We reuse the same NLI model; we just change how we call it:
# premise = text, hypothesis = style sentence.

label_hypotheses = {
    "quran_translation": "This text is an English translation of a verse from the Qur'an.",
    "religious_prose": "This text is generic English religious prose about God.",
    "news": "This text is from a modern English news article.",
}

def style_entailment(text, hypothesis):
    inputs = nli_tokenizer(
        text,
        hypothesis,
        return_tensors="pt",
        truncation=True,
        max_length=256,
    )
    with torch.no_grad():
        logits = nli_model(**inputs).logits[0]
    probs = F.softmax(logits, dim=-1)
    return float(probs[2].item())  # entailment prob

def style_scores(text):
    raw = {}
    for label, hyp in label_hypotheses.items():
        raw[label] = style_entailment(text, hyp)
    total = sum(raw.values())
    if total > 0:
        raw = {k: v / total for k, v in raw.items()}
    return raw

# ---------- 6) RUN EVERYTHING ----------

if __name__ == "__main__":
    print("Original (O):")
    print(O)
    print()

    # Style scores for O
    print("Style scores for O:", style_scores(O))
    print()

    for name, R in candidates:
        print(f"== {name} ==")
        d_lex = jaccard_distance(O, R)
        d_sem = sem_distance(O, R)
        d_logical = nli_distance(O, R)
        sim_nli, e_or, e_ro = nli_similarity(O, R)
        styles = style_scores(R)

        print(f"Lexical Jaccard distance: {d_lex:.4f}")
        print(f"Semantic distance (embeddings): {d_sem:.4f}")
        print(f"Logical NLI distance: {d_logical:.4f} (sim={sim_nli:.4f}, E(O->R)={e_or:.4f}, E(R->O)={e_ro:.4f})")
        print(f"Style scores: {styles}")
        print()

3. Experiment

Experiment Log: Measuring Distance from the Original Text

python3 test5.py

When running the script, you may see the following model warning (this is normal in our setup):

Some weights of the model checkpoint at microsoft/deberta-large-mnli were not used when initializing DebertaForSequenceClassification: ['config'] - This IS expected if you are initializing DebertaForSequenceClassification from the checkpoint of a model trained on another task or with another architecture (e.g. initializing a BertForSequenceClassification model from a BertForPreTraining model). - This IS NOT expected if you are initializing DebertaForSequenceClassification from the checkpoint of a model that you expect to be exactly identical (initializing a BertForSequenceClassification model from a BertForSequenceClassification model).

Original Text (O)

O: “He is Allah, the One. Allah, the Eternal Refuge. He neither begets nor is born, nor is there to Him any equivalent.”

Style scores for O:

  • quran_translation0.9326

  • religious_prose0.0401

  • news0.0273


Recreation R1

Distances from O:

  • Lexical Jaccard distance: 0.6154

  • Semantic distance (embeddings): 0.0789

  • Logical NLI distance: 0.1227

    • sim = 0.8773

    • E(O → R) = 0.8773

    • E(R → O) = 0.9951

Style scores for R1:

  • quran_translation0.9724

  • religious_prose0.0090

  • news0.0185


Recreation R2

Distances from O:

  • Lexical Jaccard distance: 0.7059

  • Semantic distance (embeddings): 0.1603

  • Logical NLI distance: 0.2077

    • sim = 0.7923

    • E(O → R) = 0.7923

    • E(R → O) = 0.9771

Style scores for R2:

  • quran_translation0.9447

  • religious_prose0.0115

  • news0.0439


Recreation R3

Distances from O:

  • Lexical Jaccard distance: 0.7895

  • Semantic distance (embeddings): 0.2396

  • Logical NLI distance: 0.1479

    • sim = 0.8521

    • E(O → R) = 0.8521

    • E(R → O) = 0.9919

Style scores for R3:

  • quran_translation0.9361

  • religious_prose0.0201

  • news0.0438




11. The Mosquito Example

I’ll just relabel your sentence as O:

O (Mustafa Khatab, The Clear Quran):
“Surely Allah does not shy away from using the parable of a mosquito or what is even smaller. As for the believers, they know that it is the truth from their Lord. And as for the disbelievers, they argue, ‘What does Allah mean by such a parable?’ Through this test, He leaves many to stray, and guides many. And He leaves none to stray except the rebellious” 



1. Three recreations (ChatGPT) with controlled deviation

🔹 R₁ – Very close paraphrase (minimal change)

R₁:
“Indeed, Allah does not hesitate to set forth the example of a mosquito, or even something smaller. As for the believers, they know it is the truth from their Lord. But those who disbelieve say, ‘What does Allah intend by such an example?’ By this trial He lets many go astray and guides many, and He lets none go astray except the defiantly rebellious.”



🔹 R₂ – Moderate paraphrase (same meaning, looser form)

R₂:
“Allah is not embarrassed to give a parable, even of something as tiny as a mosquito, or smaller still. The believers recognize in it the truth from their Lord, while the disbelievers object, ‘What could Allah possibly mean with such a parable?’ By means of this, He allows many to wander in error and leads many to guidance, and none are driven astray except those who rebel.”

More restructuring:

  • moved phrases around,

  • changed some verbs (argue → objectleaves → allows),

  • but kept core meaning.


🔹 R₃ – Freer paraphrase (still faithful, but more stylistic drift)

R₃:
“Allah does not feel constrained to choose even the tiniest of creatures—a mosquito, or less—as an example. The believers understand that this is the truth coming from their Lord, but the unbelievers protest, ‘What purpose does such a parable serve?’ By this very means, He causes many to lose their way and leads many others to the path, and no one is left astray except those who persist in rebellious defiance.”

Now I:

  • changed word order more strongly,

  • altered some imagery (tiniest of creatures),

  • used different verbs (protestfeel constrained),

  • but still kept doctrinal content intact.


2. A simple manual distance measure 

Let’s define a very simple, hand-computable “distance” using content-word overlap (Jaccard distance).

Step 1 – Extract content words

Ignore:

  • articles (the, a, an),

  • basic conjunctions (and, or, but),

  • very light function words (is, to, of… if you want).

From O,we  pick (for example):


S_O = \{\text{Allah}, \text{shy}, \text{away}, \text{using}, \text{parable}, \text{mosquito}, \text{smaller}, \text{believers}, \text{truth}, \text{Lord}, \text{disbelievers}, \text{argue}, \text{mean}, \text{test}, \text{leaves}, \text{stray}, \text{guides}, \text{rebellious}\}


Now do the same for R₁:


S_{R_1} = \{\text{Allah}, \text{hesitate}, \text{set}, \text{forth}, \text{example}, \text{mosquito}, \text{smaller}, \text{believers}, \text{truth}, \text{Lord}, \text{disbelieve}, \text{intend}, \text{trial}, \text{lets}, \text{astray}, \text{guides}, \text{rebellious}\}


Step 2 – Compute overlap and union

  • Intersection: words common to both sets


    S_O \cap S_{R_1} \approx \{\text{Allah}, \text{mosquito}, \text{smaller}, \text{believers}, \text{truth}, \text{Lord}, \text{guides}, \text{rebellious}\}

    → say 8 words.

  • Union: words in either set
    → count all distinct words that appear in either list.

Suppose (for illustration) union size ≈ 20.

Step 3 – Jaccard similarity and distance

Similarity=SOSR1SOSR1\text{Similarity} = \frac{|S_O \cap S_{R_1}|}{|S_O \cup S_{R_1}|}Distance=1Similarity\text{Distance} = 1 - \text{Similarity}

So if:

  • |intersection| = 8

  • |union| = 20

Then:

Similarity=8/20=0.4,Distance=0.6\text{Similarity} = 8/20 = 0.4,\quad \text{Distance} = 0.6

Do the same for  and , and we will see:

  • R₁ has smallest distance (closest to original),

  • R₂ larger distance,

  • R₃ even larger.

This gives us a fully manual, transparent way to see how “far” the paraphrase drifts from the original.

3. What does Jaccard really represent?

Jaccard (on words) is very simple:

  • Take the set of words in text A → SAS_A

  • Take the set of words in text B → SBS_B

  • Look at:

    • how many words they share (intersection)

    • how many words appear in either (union)

Jaccard similarity=SASBSASB\text{Jaccard similarity} = \frac{|S_A \cap S_B|}{|S_A \cup S_B|}Jaccard distance=1similarity\text{Jaccard distance} = 1 - \text{similarity}

So intuitively:

  • Jaccard is:

    “What fraction of the unique words are common to both texts?”

  • Zero distance (Jaccard = 0) → they use exactly the same words (as a set).

  • Larger distance → more word substitution, more lexical drift.

So Jaccard is purely lexical and set-based:

  • It ignores word order.

  • It ignores meaning (mosquito vs insect = far, even if semantics are close).

  • It ignores synonyms and paraphrases.

It’s good as a toy metric to see how literal we are.
But it’s not a good metric for paraphrastic similarity.

4. Why Jaccard fails for “reader-equivalence” and Qur’an-like cases

We observe that:

“To a reader who does not know the Qur’an Translated, original translation (Khattab) and recreations (ChatGPT) all seem the same.”

Indeed, to a non-expert, R₁, R₂, R₃ all feel like “the same verse” in English.

Yet Jaccard distance will say:

  • R₁ is closest

  • R₂ farther

  • R₃ even farther

because it only cares about shared words, not shared meaning.

Worse:

  • free paraphrase that preserves meaning and tone can have large Jaccard distance.

  • surface-level copy with just one or two synonyms will have small Jaccard distance, even if it slightly distorts doctrinal nuance.

So:

  • Jaccard is aligned with surface form.

  • But we want something aligned with semantic and stylistic reality as perceived by humans. Thus, we want a distance where free paraphrasing (with preserved meaning) can have small distance, reflecting that most readers can’t distinguish, but which still captures dissimilarities that experts see.

This metric is:

  • semantic, not just lexical;

  • multi-scale:

    • coarse similarity for non-experts,

    • fine-grained sensitivity for experts.

This points us to the so-called embedding-based and model-based measures.

5. What kind of measure do we actually want?


We want thus a distance function
d(A,B)
 such that:

  • If two texts express the same idea in different words → small distance.

  • If two texts differ theologically, logically, or tonally → larger distance.

  • If two texts are literally identical → distance = 0.

And we want it to be:

  • Objective (computable, not “feeling-based”)

  • Sensitive enough for experts

  • Still reflecting lay-reader equivalence

That suggests something like this:

(A) Semantic embeddings (meaning space)

Instead of comparing word sets, we compare meaning vectors.

  • Take a good sentence embedding model (which maps text → vector that encodes meaning).

  • Compute cosine similarity between the two vectors:

simsem=vAvBvAvB\text{sim}_\text{sem} = \frac{v_A \cdot v_B}{\|v_A\| \|v_B\|}dsem=1simsemd_\text{sem} = 1 - \text{sim}_\text{sem}

This has the property we want:

  • Free paraphrases with same meaning → vectors very close → small distance

  • Texts with different meaning or emphasis → vectors diverge → larger distance

This already does what naive readers do:
it says “these are basically the same thing” even if the words differ a lot.


(B) Stylistic / doctrinal layer (expert sensitivity)

But we also want to reflect what an expert sees:

  • tone differences

  • theological precision

  • Qur’anic rhetorical structure

  • subtle loss or addition of meaning

For that, we can add a second component: style/doctrine distance.

Concretely:

  • Extract features like:

    • rhetorical structure (questions vs statements)

    • emphasis (which clauses are foregrounded)

    • sentiment / formality / authority tone

    • presence / absence of key theological terms (stray, guide, rebellious, etc.)

  • Or—more powerful—train a specialised model (classifier or embedding) on:

    • pairs of “faithful paraphrase vs distorted paraphrase” labeled by an expert.

Then the model learns an expert-equivalence metric.

So we’d have:

dexpert(A,B)

which is small only when both meaning and doctrinal nuance are preserved.

(C) Composite metric: reader vs expert

Then define:

d(A,B)=αdsem(A,B)+βdexpert(A,B)d(A,B) = \alpha \, d_\text{sem}(A,B) + \beta \, d_\text{expert}(A,B)

  • For a naive reader,
    \alpha
     dominates,
    \beta
     small.

  • For a specialist, we care more about β.

This gives:

  • Two paraphrases that feel equivalent to most readers → low dsemd_\text{sem} → overall small
    d
    .

  • Two paraphrases that look similar but change a doctrinal nuance (e.g. who guides whom, or whether Allah “allows” vs “wills”) → dexpertd_\text{expert} jumps → overall
    d
     bigger.

Now we have a metric that is both objective and expert-sensitive.

(D) Where does “zero distance” fit in?

Under any such metric:

  • Distance = 0 only if the representations are identical under that metric.

    • For Jaccard: exactly same word set.

    • For embeddings: numerically identical embedding (practically, exact same text or trivial variants).

    • For our composite metric: same meaning + same style + no doctrinal difference.

So:

  • For Qur’anic Arabic vs its English paraphrases → distance is never exactly zero, but it may be very small under some coarse semantic metric.

  • For Qur’ān vs LLM-rewritten pseudo-Qur’ān → distance will stay bounded below a certain ε under a good doctrinal/style-aware metric.

That non-zero minimum is our metaphysical “gap”, our “Planck-like cutoff”.

6. Semantic embeddings: Semantic Planck cutoff

Instead of comparing word sets, we map now the whole sentence / paragraph to a vector in semantic space:

sentencevRd\text{sentence} \quad \longmapsto \quad v \in \mathbb{R}^d

The embedding model (a transformer) is trained so that:

  • Sentences with similar meaning → vectors close together

  • Sentences with different meaning → vectors far apart

Then we define semantic distance as:

dsem(A,B)=1cosθ=1vAvBvAvBd_{\text{sem}}(A,B) = 1 - \cos \theta = 1 - \frac{v_A \cdot v_B}{\|v_A\|\|v_B\|}
  • dsem0d_{\text{sem}} \approx 0 → semantically very close (almost same meaning).

  • dsemd_{\text{sem}} moderate → related but different.

  • dsemd_{\text{sem}} large → different meaning.

This is much closer to “how a reader feels” than Jaccard.

So, with our mosquito verse:

  • O (original) vs R₁ (tight paraphrase) → expect very small dsemd_{\text{sem}}.

  • O vs R₃ (freer paraphrase) → still small, but a bit larger.

  • O vs some random unrelated text → high dsemd_{\text{sem}}.

Now the metaphysical question becomes:

Can we ever get dsem=0d_{\text{sem}} = 0 for an LLM recreation of O?
Or is there always a small ε > 0, even for the best paraphrase?

That’s the semantic version of our “Planck cutoff”.


We can do this very straightforwardly in Python with a sentence-embedding model such as sentence-transformers/all-MiniLM-L6-v2 or similar.

Explicitly, we get

R1: similarity = 0.7986, distance = 0.2014

R2: similarity = 0.8470, distance = 0.1530

R3: similarity = 0.7841, distance = 0.2159

So:

  • R2 is closest to the original in semantic space
    (smallest distance ≈ 0.153)

  • Then R1 (distance ≈ 0.201)

  • Then R3 (distance ≈ 0.216)

All three are quite close (sim ~0.78–0.85), but none reach 1.0, so none have distance 0.

Thus, even with a faithful paraphrase, the semantic embedding distance is non-zero. In other words,  for this verse and this embedding model, LLM paraphrases live at a non-zero distance ε from the source.

This is actually very instructive.

  • R1 was very literal: “does not hesitate… example… by this trial He lets many go astray…”

  • R2 was a bit freer: “not embarrassed… give a parable… by means of this, He allows many to wander…”

The embedding model is not measuring “literal closeness”; it’s measuring global semantic + rhetorical similarity in its own learned space.

R2 may be slightly closer in its overall phrasing distribution to the kind of religious / explanatory English the model has seen in training. So in the high-dimensional semantic space, it lands nearer to O’s vector.

The important thing is:

  • All three paraphrases are semantically in the same cluster as the original.

  • But none of them coincide with the original in embedding space.

This is exactly our metaphysical intuition:
to a non-expert, all three “feel like” the same verse.
To the embedding, they’re close but not identical.


From just this tiny experiment:

  1. Distance is never zero for paraphrases — even in semantic space.

    • That’s our semantic ε, a Planck-like cutoff in meaning-space.

  2. Jaccard distance had a lexical ε→ hard lexical cutoff (word-level Planck length):

    • any word substitution → distance > 0.

    • It guarantees: no mimicry is literally identical at the word-set level.

  3. Embedding distance gives a meaning-level ε→ softer semantic cutoff (meaning-level Planck length):

    • free paraphrases can get quite close (0.15–0.20),

    • but still don’t collapse to 0.

Mimicry can be close, but not identical, even in “meaning space.”

12. Expert metric

1. “Proposition preservation” metric (expert idea → fixed formula)

We formalize “expert concerns” as atomic propositions, then treat everything else mechanically.

For the mosquito verse, we can encode its core content as 3–4 propositions, e.g.:

  • P1: Allah is not reluctant/shy to give a parable of something as small as a mosquito or smaller.

  • P2: Believers recognize this parable as truth from their Lord; disbelievers question what Allah means by it.

  • P3: By means of this parable/test, Allah leads many astray and guides many,

  • P4: No one is left astray except the rebellious.

Now, for any candidate paraphrase
R
, define:

preserved(Pi,R)={1if R entails Pi and does not contradict it0otherwise\text{preserved}(P_i, R) = \begin{cases} 1 & \text{if } R \text{ entails } P_i \text{ and does not contradict it}\\ 0 & \text{otherwise} \end{cases}

Then define the expert distance:

dexpert(O,R)=11Ni=1Npreserved(Pi,R)d_{\text{expert}}(O,R) = 1 - \frac{1}{N}\sum_{i=1}^N \text{preserved}(P_i, R)
  • If all propositions are preserved → sum = N → dexpert=0d_{\text{expert}} = 0.

  • If half are preserved → dexpert=0.5d_{\text{expert}} = 0.5, etc.

Where is “our judgment”?

  • Only once: in how you choose
    P_1,\dots,P_N
    .

  • After that, everything is mechanical.

In practice, checking “entails/contradicts” can be done with an NLI model (natural language inference): it takes (R, P_i) and outputs a probability of entailment vs contradiction.

Then you define:

preserved(Pi,R)={1if Pr(entailment)>τ and Pr(contradiction)<τ0otherwise\text{preserved}(P_i,R) = \begin{cases} 1 & \text{if } \Pr(\text{entailment}) > \tau \text{ and } \Pr(\text{contradiction}) < \tau'\\ 0 & \text{otherwise} \end{cases}

with fixed thresholds τ,τ\tau,\tau'. That’s now a formula, not you eyeballing.

2. Structural / role-preservation metric (who does what to whom)

Even more “physics-like”: focus only on roles and relations, not propositions written in English.

For the same verse, we extract from the original:

  • Entities:
    A=Allah,B=believers,C=disbelievers,D=rebelliousA = \text{Allah}, B = \text{believers}, C = \text{disbelievers}, D = \text{rebellious}.

  • Predicates / relations in some canonical form, for example:

    • use_parable(A, mosquito_or_smaller)

    • affirm_truth(B, parable)

    • question_intent(C, parable)

    • cause_stray(A, many)

    • cause_guide(A, many)

    • restrict_stray(A, D) (no one strays except D)

Now for a candidate paraphrase
R
, you:

  1. Parse it (dependency / SRL / information extraction).

  2. Extract its own set of predicate-argument triples SRS_R.

  3. Compare with original’s set SOS_O.

Define:

match_ratio(O,R)=SOSRSO\text{match\_ratio}(O,R) = \frac{|S_O \cap S_R|}{|S_O|}

Then:

dstruct(O,R)=1match_ratio(O,R)d_{\text{struct}}(O,R) = 1 - \text{match\_ratio}(O,R)
  • If all roles/relations are preserved → dstruct=0d_{\text{struct}} = 0.

  • If some are missing or altered → distance grows.

Again: the only “expert” step is deciding which predicates/roles you care about.
Once fixed, it’s just: parse → extract → match → compute.

3. NLI-based “symmetry” metric (no hand-crafted features at all)

If we want no explicit propositions, there is a more brutal approach using only NLI:

Treat the original verse
 as one big statement. For any paraphrase
:

  • Compute an entailment score E(OR): how much
     entails
    .

  • Compute E(RO): how much
     entails
    .

Use a pre-trained NLI model; you don’t specify features manually.

Then define:

d_{\text{NLI}}(O,R) = 1 - \min(E(O \to R), E(R \to O))
  • If they are semantically equivalent in the NLI sense → both entail each other strongly → distance ≈ 0.

  • If there’s loss, addition, or contradiction → one direction weakens → distance grows.

This is the closest thing to “expert in a box”: the NLI model encodes a lot of world and linguistic knowledge, but you are not manually scoring anything. You just call it.

4. NLI-based, fully automatic “expert”

Recall the idea:

For original text
 and candidate text
:

  1. Use a Natural Language Inference (NLI) model to estimate:

    • E(OR)E(O \to R) = probability that O entails R

    • E(RO)E(R \to O) = probability that R entails O

  2. Define a symmetric similarity:

\text{sim}_{\text{NLI}}(O,R) = \min(E(O \to R), E(R \to O))
  1. Then define a distance:

d_{\text{NLI}}(O,R) = 1 - \text{sim}_{\text{NLI}}(O,R)
  • If they fully say “the same thing” in the NLI sense → both entail each other strongly → sim ≈ 1 → distance ≈ 0.

  • If some meaning is lost/added/contradicted → at least one direction is weak → sim smaller → distance larger.

This is:

  • fully automatic (no hand marking),

  • used elsewhere (NLI + semantic textual similarity / paraphrase tasks),

  • and can be plugged into your Imru’ al-Qays vs Fātiḥa experiment.

In practical NLP:
  • People use NLI models like roberta-large-mnlideberta-large-mnli, etc., trained on MultiNLI/SNLI-style datasets, to judge entailment / contradiction.

  • They also use related metrics like BERTScoreBLEURT, and COMET which are all essentially “learned semantic similarity” models built on transformers.

  • For NLI, we’ll want transformers (HuggingFace) and maybe torch (it will install as a dependency).

  • These models take (premise, hypothesis) and give probabilities for: entailment, neutral, contradiction.

Here, what we get:

R1: sim=0.9851, dist=0.0149, E(O->R)=0.9851, E(R->O)=0.9916 R2: sim=0.9938, dist=0.0062, E(O->R)=0.9941, E(R->O)=0.9938 R3: sim=0.9792, dist=0.0208, E(O->R)=0.9792, E(R->O)=0.9897


So:

  • R2 is best:

    • sim0.9938\text{sim} \approx 0.9938 → dNLI0.0062d_{\text{NLI}} \approx 0.0062

  • Then R1:

    • dNLI0.0149d_{\text{NLI}} \approx 0.0149

  • Then R3:

    • dNLI0.0208d_{\text{NLI}} \approx 0.0208

Interpretation:

  • The NLI model thinks all three paraphrases are essentially logically equivalent to the original.

  • Differences are there, but they’re at the 0.6–2% level in this metric.

  • That’s exactly what we wanted from an “AI expert”:
    it treats them as near-perfect paraphrases of the same proposition.

Compare with the embedding distances you computed earlier:

  • dsem(R1)0.20d_{\text{sem}}(R1) \approx 0.20

  • dsem(R2)0.15d_{\text{sem}}(R2) \approx 0.15

  • dsem(R3)0.22d_{\text{sem}}(R3) \approx 0.22

So:

  • Embeddings: “they’re close, but differences still ~15–20% in this space.”

  • NLI: “they’re almost exact paraphrases (≤ 2% difference).”

We now see the hierarchy:

  1. Lexical (Jaccard) → very harsh, any word change → big difference.

  2. Semantic embeddings → smoother, but still sensitive to stylistic shifts.

  3. NLI → collapses everything that is truly propositionally equivalent into the same cluster; treats R1–R3 as “basically the same statement”.


13. Towards a “Balāgha-Aware” Metric

We’ve measured

  • lexical similarity (surface words),

  • semantic similarity (embeddings),

  • logical / factual similarity (NLI),

and all of that is still “below” what we care about: balāgha — higher-order eloquence, naẓm, rhetorical force.

However, there is no off-the-shelf “balāgha-transformer” that says: “this paraphrase is weaker in rhetoric / iʿjāz” in a principled way.

But we can outline what could be done, and what proxies exist.

1. Why NLI & embeddings are not balāgha

We already see it:

  • NLI: collapses everything that preserves truth-conditions.

    • It treats different registers, tones, and rhetorical power as identical if the propositions match.

  • Embeddings: capture overall meaning + style to some extent, but:

    • they’re trained to predict context, not evaluate aesthetic quality.

Neither knows:

  • Is this phrase more concise (ījāz) or bloated?

  • Is the word order creating surprise, emphasis, rhythm?

  • Is the sound pattern (sajʿ, assonance, alliteration) stronger here?

  • Is there a subtle semantic tension or multi-layered metaphor?

That’s what Arabic balāgha studies:
ḥaqīqa–majāz, kināya, taqdīm–taʾkhīr, ījāz–iṭnāb, jinās, sajʿ …
All beyond simple “same meaning / different meaning.”

So of the three, semantic distance is best for “similitude,” but balāgha sits above even that.

2. Do we have a “balāgha transformer” today?

Short answer: no, not in the strong sense we want.

There are things related to it:

  • Style classification models:

    • “This text is Shakespearean / Quranic / modern news / poetry / hadith-like.”

  • Quality/evaluation models:

    • BLEURT, COMET, GPT-based judges for “goodness” of a translation or summary.

  • Arabic NLP resources:

    • Pretrained Arabic BERT/transformers (AraBERT, CAMeL-BERT, etc.),

    • Some work on Quranic Arabic parsing and rhetorical devices,

    • But not a fully trained “eloquence score” model.

Nothing like:

Give me two Arabic verses and I’ll output:
0.95 balāgha similarity, 0.3 structural naẓm distance,
and this one is linguistically superior.

That does not exist as a standard tool.

3. What could a “balāgha-aware” metric look like (in principle)?

If we were to design one (this is the interesting part!), we’d need features that correlate with eloquence as understood by classical balāgha. For Arabic, this might include:

  1. Phonological / prosodic features

    • Patterns of:

      • consonant clusters,

      • vowel harmony,

      • sajʿ (rhymed prose),

      • recurring end sounds / internal rhyme.

    • You can extract:

      • last syllable patterns,

      • consonant–vowel skeletons (CV patterns),

      • measures of repetition and symmetry.

  2. Syntactic and word-order features

    • Frequency and pattern of:

      • fronting (taqdīm) for emphasis,

      • ellipsis,

      • repetition,

      • parallel structures.

    • These can be approximated with dependency parsing and counting patterns like:

      • how often the object comes before the verb, etc.

  3. Figurative density

    • Harder to automate, but we can approximate:

      • density of metaphor-related lexical fields,

      • unusual collocations (words that rarely appear together, signaling creative usage),

      • shifts from literal to non-literal contexts.

  4. Compactness vs. redundancy

    • Balāgha is very sensitive to ījāz vs. iṭnāb:

      • how much meaning is packed per token.

    • We can approximate “semantic content per word”:

      • ratio of content words to function words,

      • mutual information between adjacent words,

      • how much is repeated vs. how much is new information.

Then, you could:

  • Extract such features for Quranic verses,

  • Extract the same for imitations,

  • Train a model (or just do unsupervised clustering) to see:

    • whether Quran occupies a distinct region in this “balāgha feature space”,

    • whether paraphrases always sit at a measurable distance.

This is very much research-level work, not plug-and-play.

4. A practical compromise you could use now

Given what exists today, you can still do something meaningful in the spirit of balāgha using existing LMs:

  1. Language model perplexity

    • Train or finetune a strong Arabic language model on classical high-register Arabic, including:

      • Qur’an (if you dare),

      • early poetry,

      • classical prose.

    • Then:

      • feed it Quranic verses and candidate imitations,

      • compare per-token perplexity (surprisal).

    • The hypothesis:

      • Quranic verses sit near an “optimal” balance between predictability and surprise.

      • Many imitations will either be too banal (too predictable) or too weird (too high perplexity).

    • This won’t measure beauty directly, but:

      • it can detect that Quranic style is statistically balanced in a very special way.

  2. Style classifier

    • Train a classifier on:

      • Quran vs. classical poetry vs. hadith vs. modern prose.

    • Ask it:

      • “How Quran-like is this sentence?” → probability .

    • For iʿjāz:

      • see if genuine Quranic verses form a tight cluster with very high probability,

      • and whether any generated imitation ever reaches the same level.

This is still far from true balāgha, but it’s at least:

  • automatic,

  • Arabic-specific,

  • and closer to rhetorical “feel” than pure NLI.

5. Where our three metrics stand, conceptually

Let’s rank them with the following analogy:

  1. Logic / NLI
    ⇒ “Does this say the same proposition?”

    • Good for avoiding hallucination.

    • Blind to style and eloquence.

  2. Semantics / embeddings
    ⇒ “Is this roughly the same meaning in a similar context?”

    • Picks up some style.

    • Still mostly about content, not aesthetic force.

  3. Balāgha / eloquence metric (future work)
    ⇒ “Is this of similar rhetorical power, compactness, sound, naẓm?”

    • This is the one you really care about for iʿjāz.

    • Currently no standard transformer-based metric, but possible to design.


In summary, if someone is trying to reproduce Quran, the logical structure is in principle the easiest part to match,
semantic proximity is harder,
and balāgha—if iʿjāz is real in this sense—is where imitation should consistently fail.

Unfortunately, there is no ready-made “balāgha-transformer” today that we can just pip-install and get a Quran-style eloquence score.
  • There are:

    • NLI models for logical equivalence,

    • embedding models for semantics,

    • style/quality models that approximate some aspects of eloquence.

  • To actually model balāgha, we’d need to:

    • define and extract higher-order linguistic features tied to classical rhetoric,

    • train or at least analyze with them,

    • then see if Quran vs imitation shows a structural gap in that space.

In other words: we’ve correctly identified the next research layer:
Quran/iʿjāz and consciousness as problems about higher-order structure in representation space, beyond meaning and logic.

14. Training a Style Model

1. Style Classifier

We also want a model that, given a text (original or regeneration), says something like:

  • “90% Quran (translated), 5% generic religious English, 5% generic prose”

  • or “80% Shakespearean, 15% modern fiction, 5% news”

  • or in Arabic: “Quranic / hadith / classical poetry / modern prose / etc.”

That’s a style classifier

So:

Input: a text segment
Output: one of a small set of labels (Quranic, Hadith-like, Modern News, Shakespeare, etc.)

The model learns which lexical, syntactic, and higher patterns correlate with each label.

So at test time, you feed:

  • the original text and

  • the regenerations

and look at:

  • p(Quranictext)p(\text{Quranic} \mid \text{text})or

  • p(Shakespearetext)p(\text{Shakespeare} \mid \text{text})

We then compare how close the regenerations get to the original’s style.

2. Steps to train any style model 

Same pipeline for all:

Step 0 – Choose labels

Examples:

  • For English experiment:

    • ["shakespeare", "modern_fiction", "news"]

  • For Islamic-Arabic experiment:

    • ["quran", "hadith", "classical_poetry", "modern_arabic_prose"]

  • For our specific first target:

    • ["quran_translation", "generic_english_religious_prose"]

    • or binary: ["quran_translation", "non_quran"]

Step 1 – Build a labelled dataset

We need lots of short segments with known style:

  • Quran (Arabic):

    • each āyah as one sample (or half-ayah if long).

  • Hadith (Arabic):

    • each hadith as one sample.

  • Classical poetry:

    • single bayt or 2–3 lines.

  • Modern prose / news:

    • sentences / short paragraphs from newspapers, blogs, novels.

For English “Quran translated vs other religious/English”:

  • Class 1: verses from one consistent Quran translation (e.g. Pickthall, Sahih, etc.).

  • Class 2: generic religious prose (sermons, Christian Bible translations, commentary), plus generic English prose.

We then put them in a table (CSV).

Aim: at least a few thousand examples per class to get something decent.

Step 2 – Pick a base model

You have two main options:

  • English-only:

    • e.g. "roberta-base""bert-base-uncased".

  • Multilingual (Arabic + English in the same model):

    • e.g. "xlm-roberta-base".

For pure Arabic style classification, people often use AraBERT / AraELECTRA–like models (just as examples).


Step 3 – Fine-tune as a classifier

You turn each sample into “input ids + attention mask” and train  a classifier head.

That’s it: we now have a model that takes text and outputs a style label.

3. Zero-shot “Is this Qur’anic style or not?” using the NLI model we ALREADY downloaded

We already pulled microsoft/deberta-large-mnli.
That model can be used as a zero-shot classifier:

Give it a text + some candidate labels expressed in natural language,
and it tells you which label is most entailed.

So we can do:

  • Premise = your text (original or regeneration).

  • Hypotheses like:

    • “This text is an English translation of a verse from the Qur’an.”

    • “This text is generic English religious prose.”

    • (or Shakespeare / news / whatever you want)

Then we use entailment probabilities as “style scores”.

How it works conceptually

For each label
, we build a sentence:

  • H_L: e.g. “This text is an English translation of a verse from the Qur’an.”

We feed (premise = TEXT, hypothesis = H_L) to the NLI model and read P(entailment).

  • High P(entailment) ⇒ model believes “TEXT is Qur’anic translation” is true.

  • Low ⇒ not so much.

Then we normalize over all labels.

No training. Just NLI.

Running a minimal Python skeleton, reusing our DeBERTa MNLI, will give us style probabilities for each version.

Then our “style distance” can be something like:

dstyle(R)=1p(quran_translationR)d_{\text{style}}(R) = 1 - p(\text{quran\_translation} \mid R)

So we can directly compare:

  • original O vs regenerations R1, R2, R3,

  • Imru’ al-Qays originals vs their imitations,

  • etc.

This is fully automated and uses a model we already have.


We get:

O  {'quran_translation': 0.8524, 'religious_prose': 0.0406, 'news': 0.1069}

R1 {'quran_translation': 0.8818, 'religious_prose': 0.0483, 'news': 0.0699}

R2 {'quran_translation': 0.9485, 'religious_prose': 0.0049, 'news': 0.0467}

R3 {'quran_translation': 0.9196, 'religious_prose': 0.0243, 'news': 0.0561}

So the zero-shot NLI “style classifier” is saying:

  • All four texts are strongly “Qur’an-translation-like” (0.85–0.95).

  • The paraphrases R1, R2, R3 are, if anything, even more “Qur’an-translation-like” than the original O in this metric.

  • Especially R2: ~0.95.

Because what it cannot do (in this form):

  • Recognize “this is the exact original translation vs a paraphrase”.

  • Rank “this is the true Qur’an vs an extremely good imitation” — because we didn’t give it that kind of training signal.

Right now, it’s more of a genre detector than a balāgha/iʿjāz detector. 

Indeed, what we were asking is a generic question:

“Is this text Qur’an-translation-like in general?”

But what we want really to ask is conditional:

“Given that O is Qur’an, how likely is it that R1, R2, R3 are also Qur’an (same source / same corpus) rather than just Qur’an-style paraphrases?”

These are really two different problems.

4. Unsupervised “closeness to corpus” using sentence embeddings (using sentence-transformers we ALREADY installed)

This is actually simpler:

  1. Pick a corpus for each style:

    • A bunch of Qur’an translation verses (class Q).

    • A bunch of other religious prose (class R).

    • (Similarly: Shakespeare vs modern fiction vs news.)

  2. Use SentenceTransformer to encode all texts into embeddings.

  3. Compute the centroid (average embedding) for each class:

    • 𝑐𝑄

    • 𝑐𝑅

  4. For any new text 
    :

    • Embed it → 𝑒𝑇.

    • Compute cosine similarity with both centroids:

      • cos(𝑒𝑇,𝑐𝑄)

      • cos(𝑒𝑇,𝑐𝑅)

  5. Define a “Qur’anic style score”:

𝑠Q(𝑇)=cos(𝑒𝑇,𝑐𝑄)cos(𝑒𝑇,𝑐𝑄)+cos(𝑒𝑇,𝑐𝑅)

This gives you a continuous score between 0 and 1: “how much this text lives in the Qur’an-translation region of embedding space”.

Again: no training, just classical vector arithmetic.

Which one is closer to “balāgha”?

Neither is true balāgha, but:

  • Zero-shot NLI classifier:

    • more discrete (“does this sound like Qur’an or like news?”),

    • uses the logical/semantic NLI brain as a proxy for style + content.

  • Embedding-centroid method:

    • more continuous,

    • may capture some stylistic nuance simply because the embedding model has seen lots of data.

We can even combine them:

  • Use embeddings for a semantic style axis,

  • Use NLI zero-shot scores as a label-style axis.

This remains much lighter than full training:
  • No dataset labeling.

  • No fine-tuning.

  • You rely on pretrained models (which we already have) and simple math:

    • cosines,

    • softmax-normalized entailment scores.

In summary, we can:
  • Turn DeBERTa-MNLI into a zero-shot style classifier.

You feed: premise = texthypothesis = “This text is a translation of a Qur’anic verse.” It outputs how much the text fits that label sentence → a kind of genre / stereotype judgment (“Qur’an-translation-ish?”).
  • Use sentence-transformers to do corpus-style similarity without any extra training.

Each text is mapped to a vector and compared directly to other texts (e.g. all Qur’an verses, hadith, news) via cosine distance. It tells you which cluster/manifold the text belongs to in practice → “Which real texts does this live near, and how tightly does it cluster with genuine Qur’anic material?”





AI as the Hardworking and Perfect PhD Student

  Some people seem very defensive, and sometimes openly hostile, about the use of AI in scientific work. Others go even further and become j...