← All Experiments

BuddhaBERT

Active
2×2 StudyMixed ResultsWorking Report

Hypothesis

Does contemplative training material, or a set of contemplative architectural constraints, change measurable model behavior beyond vocabulary? Cross the two factors so that the reading list and the architecture can be compared separately.

Overview

Results updated September 14, 2026 · Working report, not peer reviewed

The four-model experiment is measured. It does not establish a consistent contemplative advantage. Some internal properties changed, but uncertainty, built-in mechanisms and low predictive competence limit what those changes mean.

Read the full results report →

Four models, two questions

Two models read general text and two read contemplative literature. Within each reading list, one used an unmodified architecture and one combined three contemplative interventions. The design asks separately what changes with the reading list and what changes with the architecture. Only one checkpoint was trained per combination.

What we found

  • Confidence variance remains unresolved. The pooled point estimate fell 41.5%, but its interval spans substantial increases and decreases. Lower confidence and sampling noise complicate the apparent improvement.
  • Attention is more uniform largely by construction. Entropy rose about 3.6%. The architecture starts with roughly half its attention blended toward uniform; this does not establish learned mindfulness.
  • Perturbation robustness is inconclusive. None of the six architecture comparisons survives the supplemental multiple-test correction. Only 5–9 of 150 clean targets were correct per model, making performance ratios unstable.
  • Paraphrase consistency is mixed. Similarity increased on general text and decreased on dharma text. Comparison with unrelated questions limits any claim that increased similarity means better discrimination of meaning.
  • There is a competence cost. The modified architecture has higher masked reconstruction loss and lower accuracy on both corpora.
Paraphrase similarity rises for the modified general-text model and falls for the modified dharma model; unrelated-question similarity is also high.
Same-meaning and unrelated-question similarity, with all 50 question groups below. The architecture effects point in opposite directions across corpora.

What the models could not support

The qualitative run produced 720 unreadable continuations. No blind quality ratings or dharma question-answering accuracy are claimed. The causal-attention implementation was trained with a masked reconstruction objective; correcting the generation path did not produce usable output. This limits the instrument without proving that model size caused the problem or that the architecture had no effect.

These measurements do not establish machine consciousness. The complete report distinguishes the original measurements from the supplemental September run, explains the statistical limits, and provides figures and downloadable aggregate data.

Correction and next decision

This update replaces the earlier statement that both success criteria passed. Confidence variance is unresolved; attention entropy meets the directional threshold mechanically; the question-answering criterion cannot be evaluated. Statistical testing and the qualitative run have now been completed to the limits described here.

Further training remains conditional. A next feasibility test would need a coherent training and generation objective, a fixed readability criterion and a compute cap. No larger-model or third-corpus run is reported as completed.

Methodology

  1. Train four small transformers from scratch: vanilla or modified architecture, each on general text or contemplative literature. One checkpoint per condition; no third-corpus cells were trained.
  2. Bundle attention smoothing, a confidence-variation penalty and periodic output-head resets in the modified architecture. This experiment does not isolate the three modifications from one another.
  3. Evaluate masked-token reconstruction, confidence variance and attention entropy. Report uncertainty and distinguish an imposed mechanism from a learned capacity.
  4. Complete supplemental context-perturbation and paraphrase measurements on the same checkpoints, with the protocol frozen before supplemental scoring. Compare paraphrases with unrelated questions.
  5. Attempt a qualitative comparison using 720 continuations. Report the unreadable output as an instrument limitation; do not invent blind ratings or question-answering accuracy.
  6. Retain results, code, hashes and limitations. Question-level resampling describes these fixed models; it does not replace replication across training seeds.

Status

Active

7 of 16 tracked tasks complete. The original four-model encoder measurements are complete. No consistent contemplative advantage was established. A public working report documents the mixed results and generation limitation; further training remains conditional.


← Back to all experiments