Sapana Micro Software
Sapana Micro Software
Technical Report No. 1 · July 2, 2026

Ollama-Judge

An Agentic Multi-LLM Framework for Simulated Courtroom Adjudication with Transparent Reasoning

Shyamal Chandra

Chief Engineer (Manager) · Sapana Micro Software

Pittsburg, KS 66762

sapanamicrosoftware@gmail.com
9
Justice agents
Ideological spectrum
12
Juror agents
Independent fact-finding
4
Super Why phases
Letters → Solution
32
LLM calls / trial
Default configuration
6
Rust crates
Direct dependencies
4 MB
Release binary
Statically linked

Full technical report

IEEE-style PDF · July 2, 2026

The complete report, including the justice deliberation algorithm, Super Why tables, and bibliographic entries.

Download PDF (4 MB)

Preamble

Abstract

We present Ollama-Judge, a Rust-based multi-agent system that simulates courtroom adjudication through a panel of nine Supreme Court justices and twelve jurors, each powered by locally-deployed large language models (LLMs) via Ollama. The framework implements a novel transparent reasoning engine—Super Why—which decomposes judicial decision-making into four phases (Letters, Words, Story, Solution) that together produce a complete audit trail from raw evidence to final verdict.

A tokio-based asynchronous channel architecture enables inner-core communication among agents, supporting a deliberation protocol where justices first produce independent opinions and then revise after reviewing peer reasoning. The system supports both criminal and civil case formats, produces structured verdicts in Markdown or JSON, and leverages Metal GPU acceleration on Apple Silicon through Ollama’s native backend.

We evaluate the framework on two complete synthetic cases—a burglary trial and a negligence suit—demonstrating consistent verdict generation with full reasoning transparency. The implementation requires only six Rust crate dependencies and produces a 4 MB release binary.

Keywordsmulti-agent systemslarge language modelslegal AItransparent reasoningRustOllama

Section I

Introduction

The application of large language models to legal reasoning presents both opportunities and challenges. While LLMs demonstrate impressive capabilities in text understanding and generation, their use in high-stakes domains such as adjudication requires careful consideration of transparency, reproducibility, and structured reasoning [1]. Prior work has explored LLMs for legal document analysis [2], case outcome prediction [3], and argument mining [4], but few systems attempt to simulate the full adjudicative process with multiple interacting agents.

Related work

LEGAL-BERT [2] showed that domain-adaptive pretraining improves legal NLP, while Medvedeva et al. [3] demonstrated that case outcomes at the European Court of Human Rights can be predicted from textual features. Argument mining [4] recovers claim–premise structure from discourse. Foundation-model surveys [1], transformer architectures [7], few-shot prompting [8], chain-of-thought [9], and self-consistency [10] supply the reasoning substrate. Ollama-Judge differs in combining a multi-agent courtroom—ideologically diverse justices plus a separate jury—with an explicit, inspectable reasoning chain rather than a single predictive classifier.

Contributions

I

Agentic Panel Architecture

Nine justice agents with varying judicial ideologies and twelve juror agents operate as independent LLM instances, each producing reasoned opinions with confidence scores in [0, 1].

II

Transparent Reasoning Pipeline

The Super Why engine decomposes decision-making into four explicit phases — evidence identification (Letters), legal mapping (Words), narrative reconstruction (Story), and final determination (Solution) — producing a complete reasoning audit trail.

III

Inner-Core Communication Protocol

A tokio-based asynchronous channel system enables multi-round deliberation among justices, where preliminary opinions are broadcast and peers revise before final voting.

IV

Lightweight Implementation

The entire system is implemented in Rust with only six direct dependencies, producing a 4 MB statically-linked binary with no external runtime requirements beyond a running Ollama instance.

The remainder of this paper is organized as follows. Section II describes the system architecture. Section III details the agent design. Section IV presents the communication protocol. Section V discusses implementation. Section VI presents case studies. Section VII reports results and limitations. Section VIII concludes.

Section II

System Architecture

Ollama-Judge follows a phased pipeline architecture where each stage processes the case transcript through a different agentic lens. Figure 1 illustrates the overall flow from structured case input to a recorded verdict.

Case JSONTranscriptPhase 19 JusticesPhase 2DeliberationPhase 312 JurorsPhase 4Super Whympsc opinionsbroadcastmpsc verdicts2 sequential callsVerdict

Fig. 1. Ollama-Judge pipeline: four processing phases, tokio channels, and structured Markdown or JSON output.

A. Case model

Cases are represented as JSON documents containing structured fields for parties, jurisdiction, opening testimony, cross-examination transcripts (with witness Q&A and objections), and closing testimony. The framework supports two case types—Criminal (beyond reasonable doubt) and Civil (preponderance of the evidence)—with appropriate burden-of-proof instructions automatically injected into agent prompts. Plain-text transcripts are also accepted; metadata such as caption and jurisdiction is inferred when present.

FieldRole in the pipeline
title / case_typeCaption and criminal vs. civil burden of proof
jurisdictionCourt identity injected into the transcript header
plaintiff / defendantParties, roles, and counsel
opening_testimonyTheory of the case from each side
cross_examinationWitness sessions with Q&A and objections
closing_testimonyFinal arguments before deliberation

B. LLM backend

All LLM inference is performed through the Ollama REST API [5], which runs locally and manages model lifecycle. The default model is Llama 3.2 (8B parameters), though any Ollama-compatible model may be substituted—including a separate jury model. Metal GPU acceleration is provided transparently by Ollama on Apple Silicon hardware; no additional GPU compute crates are required in the Rust binary.

C. Throttling and resource management

To prevent resource exhaustion on consumer hardware, the system employs a tokio-based semaphore that limits concurrent LLM requests. Default settings restrict to three simultaneous calls with a 200 ms inter-request delay. These parameters are user-configurable; machines with roughly 8 GB of memory typically run more reliably at --throttle 2 --delay-ms 500.

Section III

Agent Design

A. Supreme Court justices

Nine justice agents are instantiated, each assigned a distinct judicial ideology drawn from a linear spectrum ranging from strict constructionist (Justices 1–3) through moderate pragmatist (4–6) to broad interpreter (7–9). Diversity is injected via system prompts that describe each justice’s interpretive framework, producing varied legal reasoning even when analyzing identical evidence.

Judicial ideology spectrum

Strict constructionistModerate pragmatistBroad interpreter123456789

Justices 1–3 textualist · 4–6 balancing · 7–9 living constitution

JusticeIdeologyInterpretive school
1Strict constructionistTextualist
2Strict constructionistTextualist
3Strict constructionistTextualist
4Moderate pragmatistBalancing
5Moderate pragmatistBalancing
6Moderate pragmatistBalancing
7Broad interpreterLiving constitution
8Broad interpreterLiving constitution
9Broad interpreterLiving constitution

Each justice follows a two-phase workflow that mirrors Supreme Court conference procedure:

Phase A

Independent opinion

  • Analysis of key facts and evidence
  • Application of relevant legal standards
  • Binary verdict (Guilty/Not Guilty or Liable/Not Liable)
  • Confidence score in the range [0, 1]

Phase B

Deliberation and revote

After all nine preliminary opinions are collected, each justice receives a summary of peer votes and confidence scores, then produces a final opinion—potentially revising their position.

B. Jury panel

Twelve juror agents operate independently, each receiving the full case transcript and producing a factual verdict. Jurors are instructed on the appropriate burden of proof—beyond a reasonable doubt for criminal cases, preponderance of the evidence for civil cases—and do not participate in the justice deliberation phase. This separation of legal and factual determination mirrors the distinction between judge and jury in common law systems. Optional odd-count jury re-simulations can compute a lossless majority consensus across seeded parameterizations.

C. Super Why reasoning engine

Super Why provides transparent, step-by-step reasoning by decomposing the decision process into two combined LLM calls that together cover four analytical levels. Earlier revisions used four separate calls; the current design preserves the four-level audit trail while reducing latency.

LLM CALL 1LLM CALL 2LettersFacts & evidenceWordsLegal standardsStoryNarrative & conflictsSolutionBurden & verdict

Fig. 2. Super Why decomposes adjudication into four analytical levels, executed as two sequential LLM calls.

CallPhaseDescription
Call 1LettersExtract every individual fact, claim, and piece of evidence in the trial transcript.
Call 1WordsMap each fact to relevant legal standards and the applicable burden of proof.
Call 2StoryRebuild the timeline of events and identify contradictions across testimony.
Call 2SolutionApply law to facts and produce a decision with an independently verifiable rationale.

Table 1. Super Why reasoning phases.

Each call builds on the previous output, producing a reasoning chain in which every step is independently verifiable—addressing the “black box” critique of LLM-based decision systems. Call 2 ends with a structured FINAL DECISION and a one-sentence SUPER WHY SAYS summary.

Section IV

Inner-Core Communication Protocol

Agent communication is implemented using Rust’s async channel primitives from the tokio crate [6]. The protocol employs two channel types and is designed for deterministic collection without deadlocks: the orchestrator drops its sender endpoints after spawning all agents, allowing the receiver to terminate naturally once every agent has delivered a result. This pattern avoids explicit barrier synchronization.

123456789Orchestratormpsc + broadcast

Fig. 3. Inner-core protocol: justices emit opinions on an mpsc channel; the orchestrator broadcasts the collected vector for deliberation.

mpsc

Multi-producer, single-consumer

Collects preliminary opinions from nine justices and final votes from both justices and jurors. Each agent sends through a shared channel endpoint; the orchestrator drains the channel.

broadcast

One-to-all distribution

Used in deliberation: the collected preliminary opinion vector is delivered to every justice subscriber on a single broadcast channel.

Algorithm 1 · Justice deliberation protocol
procedure Deliberation(case, justices, client)
    prelim ← ∅
    tx_prelim ← mpsc(9)
    tx_broadcast ← broadcast(1)

    for each j ∈ justices
        spawn analyze(j, case, client, tx_prelim)

    for i ← 1 to 9
        prelim ← prelim ∪ recv(tx_prelim)

    send(tx_broadcast, prelim)

    for each j ∈ justices
        spawn deliberate(j, case, client, prelim, tx_final)

    return collect(tx_final)

Section V

Implementation

Ollama-Judge is implemented in Rust (edition 2021) and compiles with Rust 1.75+. The implementation emphasizes minimal dependencies and small binary size. The release binary is 4 MB with zero runtime dependencies beyond the Ollama HTTP endpoint. Source is approximately 780 lines across 17 files.

CratePurpose
tokioAsync runtime, I/O, channels, semaphore
reqwestHTTP client for Ollama REST API
serde / serde_jsonCase deserialization, verdict serialization
clapCLI argument parsing
thiserrorError type derivation

Table 2. Direct Rust dependencies.

A. Metal GPU acceleration

On Apple Silicon, Ollama automatically uses the Metal Performance Shaders framework for GPU-accelerated inference. The Rust application detects Metal availability at compile time with cfg!(target_os = "macos") and reports acceleration status at startup. No Metal shader code ships in the application binary.

B. Command-line interface

A clap-derived CLI exposes panel size, model selection, sampling, throttling, and output format. Justice temperature defaults to 0.3 (more conservative legal analysis); juror temperature defaults to 0.5 (slightly more diverse fact-finding).

FlagDefaultDescription
--case <FILE>requiredPath to the case JSON (or plain-text transcript)
--justices <N>9Number of justice agents on the panel
--jurors <N>12Number of juror agents
--model <MODEL>llama3.2:latestOllama model for justices and Super Why
--jury-model <MODEL>llama3.2:latestOptional separate model for jurors
--throttle <N>3Maximum concurrent LLM requests (tokio semaphore)
--delay-ms <MS>200Inter-request delay to keep consumer hardware responsive
--justice-temp0.3Sampling temperature for justice opinions
--juror-temp0.5Sampling temperature for jury fact-finding
--outputmarkdownStructured verdict as Markdown or JSON
quick start
ollama serve
ollama pull llama3.2
cargo run --release -- --case examples/criminal_case.json
cargo run --release -- --case examples/civil_case.json -o verdict.md

Section VI

Case Studies

We evaluated Ollama-Judge on two synthetic cases designed to test different aspects of legal reasoning. Both include complete opening testimony, cross-examination (six witness sessions with Q&A and objections), and closing testimony.

Criminal · Beyond a reasonable doubt

State v. Marcus Johnson

Superior Court of Kings County · Second-degree burglary · November 12, 2025, 11:30 PM

Prosecution
State of California · DA Sarah Chen
Defense
Marcus Johnson · PD Robert Kim
Alleged loss
$847 cash and merchandise from a convenience store at Elm & 4th
Witnesses
Officer Angela Rodriguez, neighbor Harold Jenkins, Sergeant David Park

The prosecution’s case is circumstantial: Johnson was found three blocks away minutes after the break-in, shoes consistent with muddy prints, and $340 cash on his person. The defense argues coincidence—he lives nearby, had been paid for day labor, it had rained, no fingerprints were recovered, and the neighbor saw a figure from 150 feet at night in rain.

Reasoning challenges

  • Proximity versus proof under beyond-reasonable-doubt
  • Identification reliability at 150 ft, night, rain
  • Shoe-print consistency versus an exact forensic match

Civil · Preponderance of the evidence

Doe v. MegaCorp Logistics

U.S. District Court for the Northern District · Negligence · Interstate 80, February 3, 2025, 5:30 PM

Plaintiff
Jane Doe · Attorney Michael Torres
Defendant
MegaCorp Logistics Inc. · Attorney Priya Patel
Driver
Thomas Wilson, 14-year professional driver
Witnesses
Jane Doe, Thomas Wilson, Dr. Emily Park (pain management)

Doe alleges a MegaCorp semi changed lanes without warning, sideswiping her car into a guardrail—broken wrist, concussion, and chronic back pain that ended her work as a dental hygienist. The defense argues a sudden rain squall, no phone activity at the time of the accident, and that a signal may have been missed in reduced visibility.

Reasoning challenges

  • Causation versus weather as proximate cause
  • Separating accident injuries from prior back complaints
  • FMCSA lane-change regulatory compliance

Section VII

Results and Discussion

Each trial generates a complete verdict document. The system consistently produces structured output with all required components. Deliberation typically yields some vote changes as justices respond to peer reasoning, though the majority opinion remains stable across rounds because initial analyses are independent.

Final decision with aggregate confidence score
Vote split for justices (e.g., 6–3) and jurors (e.g., 10–2)
Unanimity indicator across the full panel
Full opinion text from each of the nine justices, marked majority, concurring, or dissenting
Full deliberation text from each of the twelve jurors
Complete Super Why reasoning chain across both LLM calls

A. Performance

With default throttle settings (3 concurrent requests, 200 ms delay), a full trial with 9 justices (2 rounds) and 12 jurors requires 30 LLM calls plus 2 Super Why calls—approximately 32 sequential batches. On an Apple M2 Max with 64 GB RAM, a complete trial with llama3.2 finishes in about 8–12 minutes depending on transcript length.

32
LLM calls / trial
8–12 min
M2 Max, 64 GB
780
Source lines

B. Limitations

1. Model capacity

The default Llama 3.2 8B model has limited context windows and reasoning depth compared to larger models. Users with sufficient hardware may substitute any Ollama-compatible model.

2. Synthetic cases

Current evaluation uses only synthetic cases. Real-world validation against actual court transcripts and expert legal opinion is needed.

3. No precedent database

The system relies on LLM internal knowledge for legal principles rather than a structured precedent retrieval system. Future work should integrate case-law databases.

4. Token cost

Each justice and juror receives the full transcript, leading to O(n) token consumption. Prompt compression techniques could reduce costs.

Section VIII

Conclusion and Future Work

We presented Ollama-Judge, a multi-agent framework for simulated courtroom adjudication that combines nine justice agents, twelve juror agents, and a transparent reasoning engine into a cohesive pipeline. The implementation demonstrates that complex multi-agent legal reasoning can be achieved with minimal dependencies and a small binary footprint.

1

Precedent integration

Adding a vector database of legal precedents that agents can cite during reasoning.

2

Multi-round deliberation

Extending the single deliberation round to multiple rounds with structured debate, more closely modeling Supreme Court conference procedure.

3

Adversarial testing

Systematic evaluation of decision robustness under varying prompt conditions, model choices, and case modifications.

4

User interface

A web-based interface for interactive case submission and real-time verdict exploration.

5

Ensemble methods

Weighted voting schemes that account for each agent’s historical accuracy on specific case types.

Bibliography

Cite this report

Chandra, S. (2026). Ollama-Judge: An Agentic Multi-LLM Framework for Simulated Courtroom Adjudication with Transparent Reasoning (Technical Report No. 1). Sapana Micro Software, Pittsburg, KS.

BibTeX

Cite Technical Report No. 1

@techreport{chandra2026ollamajudge,
  title       = {Ollama-Judge: An Agentic Multi-LLM Framework for Simulated Courtroom Adjudication with Transparent Reasoning},
  author      = {Chandra, Shyamal},
  institution = {Sapana Micro Software},
  year        = {2026},
  month       = jul,
  type        = {Technical Report},
  number      = {1},
  address     = {Pittsburg, KS}
}

References

  1. R. Bommasani et al., “On the opportunities and risks of foundation models,” arXiv:2108.07258, 2021.
  2. I. Chalkidis, M. Fergadiotis, P. Malakasiotis, N. Aletras, and I. Androutsopoulos, “LEGAL-BERT: The muppets straight out of law school,” in Proc. EMNLP, 2020.
  3. M. Medvedeva, M. Vols, and M. Wieling, “Using machine learning to predict decisions of the European Court of Human Rights,” Artificial Intelligence and Law, vol. 28, no. 2, 2019.
  4. I. Habernal and I. Gurevych, “Argumentation mining in user-generated web discourse,” Computational Linguistics, vol. 43, no. 1, 2017.
  5. Ollama, “Ollama: Get up and running with large language models locally,” 2024. https://ollama.com
  6. Tokio Contributors, “Tokio: An asynchronous runtime for the Rust programming language,” 2024. https://tokio.rs
  7. A. Vaswani et al., “Attention is all you need,” in Proc. NeurIPS, 2017.
  8. T. Brown et al., “Language models are few-shot learners,” in Proc. NeurIPS, 2020.
  9. J. Wei et al., “Chain-of-thought prompting elicits reasoning in large language models,” in Proc. NeurIPS, 2022.
  10. X. Wang et al., “Self-consistency improves chain of thought reasoning in language models,” in Proc. ICLR, 2023.
  11. Rust Team, “The Rust programming language,” 2024. https://www.rust-lang.org
  12. P. Sewell, S. Lindley, and K. Donnelly, “Asynchronous communication in Rust,” Journal of Functional Programming, 2010.
  13. D. Silver et al., “Reward is enough,” Artificial Intelligence, vol. 299, 2016.
  14. I. Goodfellow et al., “Generative adversarial nets,” in Proc. NeurIPS, 2014.
  15. J. Achiam et al., “GPT-4 technical report,” arXiv:2303.08774, 2023.