Preamble
Abstract
We present Ollama-Judge, a Rust-based multi-agent system that simulates courtroom adjudication through a panel of nine Supreme Court justices and twelve jurors, each powered by locally-deployed large language models (LLMs) via Ollama. The framework implements a novel transparent reasoning engine—Super Why—which decomposes judicial decision-making into four phases (Letters, Words, Story, Solution) that together produce a complete audit trail from raw evidence to final verdict.
A tokio-based asynchronous channel architecture enables inner-core communication among agents, supporting a deliberation protocol where justices first produce independent opinions and then revise after reviewing peer reasoning. The system supports both criminal and civil case formats, produces structured verdicts in Markdown or JSON, and leverages Metal GPU acceleration on Apple Silicon through Ollama’s native backend.
We evaluate the framework on two complete synthetic cases—a burglary trial and a negligence suit—demonstrating consistent verdict generation with full reasoning transparency. The implementation requires only six Rust crate dependencies and produces a 4 MB release binary.
Section I
Introduction
The application of large language models to legal reasoning presents both opportunities and challenges. While LLMs demonstrate impressive capabilities in text understanding and generation, their use in high-stakes domains such as adjudication requires careful consideration of transparency, reproducibility, and structured reasoning [1]. Prior work has explored LLMs for legal document analysis [2], case outcome prediction [3], and argument mining [4], but few systems attempt to simulate the full adjudicative process with multiple interacting agents.
Related work
LEGAL-BERT [2] showed that domain-adaptive pretraining improves legal NLP, while Medvedeva et al. [3] demonstrated that case outcomes at the European Court of Human Rights can be predicted from textual features. Argument mining [4] recovers claim–premise structure from discourse. Foundation-model surveys [1], transformer architectures [7], few-shot prompting [8], chain-of-thought [9], and self-consistency [10] supply the reasoning substrate. Ollama-Judge differs in combining a multi-agent courtroom—ideologically diverse justices plus a separate jury—with an explicit, inspectable reasoning chain rather than a single predictive classifier.
Contributions
Agentic Panel Architecture
Nine justice agents with varying judicial ideologies and twelve juror agents operate as independent LLM instances, each producing reasoned opinions with confidence scores in [0, 1].
Transparent Reasoning Pipeline
The Super Why engine decomposes decision-making into four explicit phases — evidence identification (Letters), legal mapping (Words), narrative reconstruction (Story), and final determination (Solution) — producing a complete reasoning audit trail.
Inner-Core Communication Protocol
A tokio-based asynchronous channel system enables multi-round deliberation among justices, where preliminary opinions are broadcast and peers revise before final voting.
Lightweight Implementation
The entire system is implemented in Rust with only six direct dependencies, producing a 4 MB statically-linked binary with no external runtime requirements beyond a running Ollama instance.
The remainder of this paper is organized as follows. Section II describes the system architecture. Section III details the agent design. Section IV presents the communication protocol. Section V discusses implementation. Section VI presents case studies. Section VII reports results and limitations. Section VIII concludes.
Section II
System Architecture
Ollama-Judge follows a phased pipeline architecture where each stage processes the case transcript through a different agentic lens. Figure 1 illustrates the overall flow from structured case input to a recorded verdict.
Fig. 1. Ollama-Judge pipeline: four processing phases, tokio channels, and structured Markdown or JSON output.
A. Case model
Cases are represented as JSON documents containing structured fields for parties, jurisdiction, opening testimony, cross-examination transcripts (with witness Q&A and objections), and closing testimony. The framework supports two case types—Criminal (beyond reasonable doubt) and Civil (preponderance of the evidence)—with appropriate burden-of-proof instructions automatically injected into agent prompts. Plain-text transcripts are also accepted; metadata such as caption and jurisdiction is inferred when present.
| Field | Role in the pipeline |
|---|---|
| title / case_type | Caption and criminal vs. civil burden of proof |
| jurisdiction | Court identity injected into the transcript header |
| plaintiff / defendant | Parties, roles, and counsel |
| opening_testimony | Theory of the case from each side |
| cross_examination | Witness sessions with Q&A and objections |
| closing_testimony | Final arguments before deliberation |
B. LLM backend
All LLM inference is performed through the Ollama REST API [5], which runs locally and manages model lifecycle. The default model is Llama 3.2 (8B parameters), though any Ollama-compatible model may be substituted—including a separate jury model. Metal GPU acceleration is provided transparently by Ollama on Apple Silicon hardware; no additional GPU compute crates are required in the Rust binary.
C. Throttling and resource management
To prevent resource exhaustion on consumer hardware, the system employs a tokio-based semaphore that limits concurrent LLM requests. Default settings restrict to three simultaneous calls with a 200 ms inter-request delay. These parameters are user-configurable; machines with roughly 8 GB of memory typically run more reliably at --throttle 2 --delay-ms 500.
Section III
Agent Design
A. Supreme Court justices
Nine justice agents are instantiated, each assigned a distinct judicial ideology drawn from a linear spectrum ranging from strict constructionist (Justices 1–3) through moderate pragmatist (4–6) to broad interpreter (7–9). Diversity is injected via system prompts that describe each justice’s interpretive framework, producing varied legal reasoning even when analyzing identical evidence.
Judicial ideology spectrum
Justices 1–3 textualist · 4–6 balancing · 7–9 living constitution
| Justice | Ideology | Interpretive school |
|---|---|---|
| 1 | Strict constructionist | Textualist |
| 2 | Strict constructionist | Textualist |
| 3 | Strict constructionist | Textualist |
| 4 | Moderate pragmatist | Balancing |
| 5 | Moderate pragmatist | Balancing |
| 6 | Moderate pragmatist | Balancing |
| 7 | Broad interpreter | Living constitution |
| 8 | Broad interpreter | Living constitution |
| 9 | Broad interpreter | Living constitution |
Each justice follows a two-phase workflow that mirrors Supreme Court conference procedure:
Phase A
Independent opinion
- Analysis of key facts and evidence
- Application of relevant legal standards
- Binary verdict (Guilty/Not Guilty or Liable/Not Liable)
- Confidence score in the range [0, 1]
Phase B
Deliberation and revote
After all nine preliminary opinions are collected, each justice receives a summary of peer votes and confidence scores, then produces a final opinion—potentially revising their position.
B. Jury panel
Twelve juror agents operate independently, each receiving the full case transcript and producing a factual verdict. Jurors are instructed on the appropriate burden of proof—beyond a reasonable doubt for criminal cases, preponderance of the evidence for civil cases—and do not participate in the justice deliberation phase. This separation of legal and factual determination mirrors the distinction between judge and jury in common law systems. Optional odd-count jury re-simulations can compute a lossless majority consensus across seeded parameterizations.
C. Super Why reasoning engine
Super Why provides transparent, step-by-step reasoning by decomposing the decision process into two combined LLM calls that together cover four analytical levels. Earlier revisions used four separate calls; the current design preserves the four-level audit trail while reducing latency.
Fig. 2. Super Why decomposes adjudication into four analytical levels, executed as two sequential LLM calls.
| Call | Phase | Description |
|---|---|---|
| Call 1 | Letters | Extract every individual fact, claim, and piece of evidence in the trial transcript. |
| Call 1 | Words | Map each fact to relevant legal standards and the applicable burden of proof. |
| Call 2 | Story | Rebuild the timeline of events and identify contradictions across testimony. |
| Call 2 | Solution | Apply law to facts and produce a decision with an independently verifiable rationale. |
Table 1. Super Why reasoning phases.
Each call builds on the previous output, producing a reasoning chain in which every step is independently verifiable—addressing the “black box” critique of LLM-based decision systems. Call 2 ends with a structured FINAL DECISION and a one-sentence SUPER WHY SAYS summary.
Section IV
Inner-Core Communication Protocol
Agent communication is implemented using Rust’s async channel primitives from the tokio crate [6]. The protocol employs two channel types and is designed for deterministic collection without deadlocks: the orchestrator drops its sender endpoints after spawning all agents, allowing the receiver to terminate naturally once every agent has delivered a result. This pattern avoids explicit barrier synchronization.
Fig. 3. Inner-core protocol: justices emit opinions on an mpsc channel; the orchestrator broadcasts the collected vector for deliberation.
mpsc
Multi-producer, single-consumer
Collects preliminary opinions from nine justices and final votes from both justices and jurors. Each agent sends through a shared channel endpoint; the orchestrator drains the channel.
broadcast
One-to-all distribution
Used in deliberation: the collected preliminary opinion vector is delivered to every justice subscriber on a single broadcast channel.
procedure Deliberation(case, justices, client)
prelim ← ∅
tx_prelim ← mpsc(9)
tx_broadcast ← broadcast(1)
for each j ∈ justices
spawn analyze(j, case, client, tx_prelim)
for i ← 1 to 9
prelim ← prelim ∪ recv(tx_prelim)
send(tx_broadcast, prelim)
for each j ∈ justices
spawn deliberate(j, case, client, prelim, tx_final)
return collect(tx_final)Section V
Implementation
Ollama-Judge is implemented in Rust (edition 2021) and compiles with Rust 1.75+. The implementation emphasizes minimal dependencies and small binary size. The release binary is 4 MB with zero runtime dependencies beyond the Ollama HTTP endpoint. Source is approximately 780 lines across 17 files.
| Crate | Purpose |
|---|---|
| tokio | Async runtime, I/O, channels, semaphore |
| reqwest | HTTP client for Ollama REST API |
| serde / serde_json | Case deserialization, verdict serialization |
| clap | CLI argument parsing |
| thiserror | Error type derivation |
Table 2. Direct Rust dependencies.
A. Metal GPU acceleration
On Apple Silicon, Ollama automatically uses the Metal Performance Shaders framework for GPU-accelerated inference. The Rust application detects Metal availability at compile time with cfg!(target_os = "macos") and reports acceleration status at startup. No Metal shader code ships in the application binary.
B. Command-line interface
A clap-derived CLI exposes panel size, model selection, sampling, throttling, and output format. Justice temperature defaults to 0.3 (more conservative legal analysis); juror temperature defaults to 0.5 (slightly more diverse fact-finding).
| Flag | Default | Description |
|---|---|---|
| --case <FILE> | required | Path to the case JSON (or plain-text transcript) |
| --justices <N> | 9 | Number of justice agents on the panel |
| --jurors <N> | 12 | Number of juror agents |
| --model <MODEL> | llama3.2:latest | Ollama model for justices and Super Why |
| --jury-model <MODEL> | llama3.2:latest | Optional separate model for jurors |
| --throttle <N> | 3 | Maximum concurrent LLM requests (tokio semaphore) |
| --delay-ms <MS> | 200 | Inter-request delay to keep consumer hardware responsive |
| --justice-temp | 0.3 | Sampling temperature for justice opinions |
| --juror-temp | 0.5 | Sampling temperature for jury fact-finding |
| --output | markdown | Structured verdict as Markdown or JSON |
ollama serve ollama pull llama3.2 cargo run --release -- --case examples/criminal_case.json cargo run --release -- --case examples/civil_case.json -o verdict.md
Section VI
Case Studies
We evaluated Ollama-Judge on two synthetic cases designed to test different aspects of legal reasoning. Both include complete opening testimony, cross-examination (six witness sessions with Q&A and objections), and closing testimony.
Criminal · Beyond a reasonable doubt
State v. Marcus Johnson
Superior Court of Kings County · Second-degree burglary · November 12, 2025, 11:30 PM
- Prosecution
- State of California · DA Sarah Chen
- Defense
- Marcus Johnson · PD Robert Kim
- Alleged loss
- $847 cash and merchandise from a convenience store at Elm & 4th
- Witnesses
- Officer Angela Rodriguez, neighbor Harold Jenkins, Sergeant David Park
The prosecution’s case is circumstantial: Johnson was found three blocks away minutes after the break-in, shoes consistent with muddy prints, and $340 cash on his person. The defense argues coincidence—he lives nearby, had been paid for day labor, it had rained, no fingerprints were recovered, and the neighbor saw a figure from 150 feet at night in rain.
Reasoning challenges
- Proximity versus proof under beyond-reasonable-doubt
- Identification reliability at 150 ft, night, rain
- Shoe-print consistency versus an exact forensic match
Civil · Preponderance of the evidence
Doe v. MegaCorp Logistics
U.S. District Court for the Northern District · Negligence · Interstate 80, February 3, 2025, 5:30 PM
- Plaintiff
- Jane Doe · Attorney Michael Torres
- Defendant
- MegaCorp Logistics Inc. · Attorney Priya Patel
- Driver
- Thomas Wilson, 14-year professional driver
- Witnesses
- Jane Doe, Thomas Wilson, Dr. Emily Park (pain management)
Doe alleges a MegaCorp semi changed lanes without warning, sideswiping her car into a guardrail—broken wrist, concussion, and chronic back pain that ended her work as a dental hygienist. The defense argues a sudden rain squall, no phone activity at the time of the accident, and that a signal may have been missed in reduced visibility.
Reasoning challenges
- Causation versus weather as proximate cause
- Separating accident injuries from prior back complaints
- FMCSA lane-change regulatory compliance
Section VII
Results and Discussion
Each trial generates a complete verdict document. The system consistently produces structured output with all required components. Deliberation typically yields some vote changes as justices respond to peer reasoning, though the majority opinion remains stable across rounds because initial analyses are independent.
A. Performance
With default throttle settings (3 concurrent requests, 200 ms delay), a full trial with 9 justices (2 rounds) and 12 jurors requires 30 LLM calls plus 2 Super Why calls—approximately 32 sequential batches. On an Apple M2 Max with 64 GB RAM, a complete trial with llama3.2 finishes in about 8–12 minutes depending on transcript length.
B. Limitations
1. Model capacity
The default Llama 3.2 8B model has limited context windows and reasoning depth compared to larger models. Users with sufficient hardware may substitute any Ollama-compatible model.
2. Synthetic cases
Current evaluation uses only synthetic cases. Real-world validation against actual court transcripts and expert legal opinion is needed.
3. No precedent database
The system relies on LLM internal knowledge for legal principles rather than a structured precedent retrieval system. Future work should integrate case-law databases.
4. Token cost
Each justice and juror receives the full transcript, leading to O(n) token consumption. Prompt compression techniques could reduce costs.
Section VIII
Conclusion and Future Work
We presented Ollama-Judge, a multi-agent framework for simulated courtroom adjudication that combines nine justice agents, twelve juror agents, and a transparent reasoning engine into a cohesive pipeline. The implementation demonstrates that complex multi-agent legal reasoning can be achieved with minimal dependencies and a small binary footprint.
Precedent integration
Adding a vector database of legal precedents that agents can cite during reasoning.
Multi-round deliberation
Extending the single deliberation round to multiple rounds with structured debate, more closely modeling Supreme Court conference procedure.
Adversarial testing
Systematic evaluation of decision robustness under varying prompt conditions, model choices, and case modifications.
User interface
A web-based interface for interactive case submission and real-time verdict exploration.
Ensemble methods
Weighted voting schemes that account for each agent’s historical accuracy on specific case types.
Bibliography
Cite this report
Chandra, S. (2026). Ollama-Judge: An Agentic Multi-LLM Framework for Simulated Courtroom Adjudication with Transparent Reasoning (Technical Report No. 1). Sapana Micro Software, Pittsburg, KS.
BibTeX
Cite Technical Report No. 1
@techreport{chandra2026ollamajudge,
title = {Ollama-Judge: An Agentic Multi-LLM Framework for Simulated Courtroom Adjudication with Transparent Reasoning},
author = {Chandra, Shyamal},
institution = {Sapana Micro Software},
year = {2026},
month = jul,
type = {Technical Report},
number = {1},
address = {Pittsburg, KS}
}References
- R. Bommasani et al., “On the opportunities and risks of foundation models,” arXiv:2108.07258, 2021.
- I. Chalkidis, M. Fergadiotis, P. Malakasiotis, N. Aletras, and I. Androutsopoulos, “LEGAL-BERT: The muppets straight out of law school,” in Proc. EMNLP, 2020.
- M. Medvedeva, M. Vols, and M. Wieling, “Using machine learning to predict decisions of the European Court of Human Rights,” Artificial Intelligence and Law, vol. 28, no. 2, 2019.
- I. Habernal and I. Gurevych, “Argumentation mining in user-generated web discourse,” Computational Linguistics, vol. 43, no. 1, 2017.
- Ollama, “Ollama: Get up and running with large language models locally,” 2024. https://ollama.com
- Tokio Contributors, “Tokio: An asynchronous runtime for the Rust programming language,” 2024. https://tokio.rs
- A. Vaswani et al., “Attention is all you need,” in Proc. NeurIPS, 2017.
- T. Brown et al., “Language models are few-shot learners,” in Proc. NeurIPS, 2020.
- J. Wei et al., “Chain-of-thought prompting elicits reasoning in large language models,” in Proc. NeurIPS, 2022.
- X. Wang et al., “Self-consistency improves chain of thought reasoning in language models,” in Proc. ICLR, 2023.
- Rust Team, “The Rust programming language,” 2024. https://www.rust-lang.org
- P. Sewell, S. Lindley, and K. Donnelly, “Asynchronous communication in Rust,” Journal of Functional Programming, 2010.
- D. Silver et al., “Reward is enough,” Artificial Intelligence, vol. 299, 2016.
- I. Goodfellow et al., “Generative adversarial nets,” in Proc. NeurIPS, 2014.
- J. Achiam et al., “GPT-4 technical report,” arXiv:2303.08774, 2023.
