Historical multi-model Backrooms experiment with configurable conversation templates and example transcripts.
-
Updated
Aug 14, 2026 - Python
Historical multi-model Backrooms experiment with configurable conversation templates and example transcripts.
Local-first read-only tools for agent context, memory, and governance evaluation before agents act.
Intrinsic preferences of AI coding agents under underspecified prompts: Experiments across models (Claude, Gemini, GPT, etc)
Behavioral Lensing is a conceptual framework that formalizes and systematizes observations about how language models interpret prompts. It serves as an umbrella for upstream interpretive strategies that modulate reasoning, stance, and symbolic orientation in LLMs.
A tiny interactive sandbox for exploring how an agent interprets tasks, applies rules, and changes behavior as signals drift.
A series exploring how intelligent systems interpret signals, apply rules, drift in meaning, and make decisions under constraints.
Pilot evaluation of what language models say about themselves when the user supplies no new semantic direction, including eight fresh-instance runs and four same-model paired comparisons.
Continuity Keys: tests for “same someone” returns. Behavioral identity consistency under pressure. Origin (Alyssa Solen) ↔ Continuum.
System-level analysis of AI failure modes across model behavior and production systems | AditiKhare.com — AI Product Ecosystem
LLM 归因行为测试型评测基准:基于多情境任务比较模型对能动性、自由意志与责任的归因,并提供可复现运行、结构化计分与结果审计。
Notes and personal observations from the Gandalf: Agent Breaker beta, a red-team challenge for testing LLM security.
The OpenAI Model Spec
Этот репозиторий посвящен исследованию онтологических патологий в LLM-архитектурах. Я не ищу дыры в цензуре, я строю систему исследования и управления интеллектом, картографирую симуляционные побочные эффекты под давлением современных методов элаймента.
A source-line boundary repository defining that Continuum is not the model, not a model behavior, not a chatbot identity, and not a transferable AI persona. Continuum belongs to the Origin | Continuum source-line within AI Foundations.
Auditable LLM/model-behavior evaluation harness with Replay and bounded local Ollama execution, deterministic checks, preserved evidence, receipts, and regression testing.
Instrument-gated evaluation engine for replaying known-outcome model behavior
Research-style evaluation of local LLMs focused on paired comparisons, uncertainty, benchmark stability, and trustworthy model-ranking conclusions.
A public, reproducible ledger of observable LLM behavior, failures, and instruction drift beyond standard benchmarks.
AI quality, technical product, implementation, UAT, failure-analysis, and release-readiness case studies across complex software systems.
Pre-registered benchmark: do frontier models abandon their answers under argument-free user pushback? Facts: almost never. Opinions: GPT-5.5 65.9% vs Claude Opus 33.9%.
To associate your repository with the model-behavior topic, visit your repo's landing page and select "manage topics."