Research
Muhammad Zane · AI safety, alignment, and interpretability.
Research
| 2026 | Working paper |
Do internal signatures predict how models play?
2026 · Working paper
|
LMI |
| 2026 | Proposed method |
What evidence would show a model represents evaluation as an internal state?
2026 · Proposed method
|
SiteLMI |
| 2026 | Working paper |
Separating persistence caused by a model from persistence caused by the world around it, with falsification criteria.
2026 · Working paper
|
SiteLMI |
Essays
| 2026 | Essay |
Does the present exist in the universe?
2026 · Essay
|
Read |
| 2026 | Essay |
How lenders might price debt collateralised by GPU fleets.
2026 · Essay
|
Read |
| 2026 | Essay |
What a functioning forward market for GPU compute would look like.
2026 · Essay
|
Read |
Instruments
| 2026 | Interactive |
A working GPT-2 in the browser: watch computation move through embeddings, attention, MLPs, and the softmax.
2026 · Interactive
|
Launch |
| 2026 | Field map |
The literature as one navigable document: dependency graph, reading pathway, runnable protocols, open problems.
2026 · Field map
|
Open |
In preparation
| 2026 | Safety | Safeguard Durability in Open Weight Models |
|
| 2026 | Cognition | Preemptive Detection of Agentic Misalignment, and Its Shelf Life |
|
| 2026 | Safety | Biosafety Red Teaming of Evo 2 |
|
| 2026 | Interp | Biological Theories of Consciousness, Mapped onto Language Model Mechanisms |
|
| 2026 | Cognition | Distress, Flourishing, and Valence in Model Output |
|
| 2026 | Cognition | Introspection and the Reliability of Model Self Report |
Collaboration
I work on safety, alignment, evaluation awareness, and interpretability, and I am glad to collaborate. Write to m@latentmindsinstitute.com.