Arcology Engine

LLM Risk Classifier

Contents

Introduction

The notes tools in llm/tools.org gate their output through this classifier: an LLM-driven risk scorer judging note text against five categories — personal health info, private/intimate info, financial info, employer-owned data, politically sensitive data. This is meant to be simple enough that it can run on a local GPU to prevent sensitive information from being sent to a remote inference server.

The runtime is Ollama over HTTP or other similar OpenAI-ish API.

"Fail closed" is the important part of this, better to be safe than sorry, or what's the point:

  • Endpoint unconfigured, unreachable, timed out, or garbage-responding → ClassificationResult.Unavailable → the calling tool redacts content. Metadata still flows; text does not.

  • --no-risk-gate ([trustLocal]) is the single explicit escape hatch for trusted local runs.

  • Verdicts cache in llm_verdict_cache (Tools.sq), keyed by content hash + model, so unchanged notes don't re-burn GPU time.

Everything here shares the computer.whatthefuck.arcology.llm package with the tools; the tangle target is a single RiskClassifier.kt.

RiskGate — cache + threshold + fail-closed wiring

The gate is what tools actually call. It wraps the classifier with the verdict cache and returns RiskGateResult whose gated flag tells the tool to withhold content. =gate(..., verbose=true)= attaches a ClassifierTranscript to the RiskBlock containing the exact request (system + user prompt with the text) and the endpoint's raw response. Cache hits carry a transcript with =rawResponse=null= (no live exchange happened); a missing classifier endpoint carries only the input. Non-verbose calls never emit transcripts, so normal envelopes stay small.

The Gate

kotlin#+name: classifier-gate
data class RiskGateResult(
    val block: RiskBlock?,
    val gated: Boolean
)

/**
 * @param classifier null when no endpoint configured; combined with
 * [trustLocal]=false this fails closed — content redacted, metadata flows.
 * [trustLocal] corresponds to the tools' --no-risk-gate flag. =--verbose=
 * ride-along behavior is per-call (see [gate]).
 */
class RiskGate(
    private val repo: ToolQueryRepository,
    private val classifier: RiskClassifier?,
    private val model: String,
    private val trustLocal: Boolean
) {
    /**
     * [verbose] bypasses the verdict cache entirely — no read, no write: a
     * verbose run always performs the live exchange the caller asked to
     * inspect. It also attaches a [[ClassifierTranscript]] to the returned
     * block (see [gate]).
     */
    suspend fun gate(text: String, threshold: Double, verbose: Boolean = false): RiskGateResult {
        if (trustLocal) return RiskGateResult(block = null, gated = false)
        if (text.isBlank()) return RiskGateResult(block = null, gated = false)

        // Verbose runs skip the cache on purpose: the point of --verbose is
        // to observe the live classifier exchange.
        val contentHash = if (verbose) null else HashUtils.sha256(text)

        // Cache hit: a prior verdict for the same text with the same model.
        // Verbose cache hits still carry a transcript — just with
        // rawResponse=null, since no live exchange happens.
        if (contentHash != null) {
            repo.selectVerdict(contentHash, model)?.let { cached ->
                val scores = RiskScores(
                    health = cached.health, privateScore = cached.privateScore,
                    financial = cached.financial, employer = cached.employer,
                    political = cached.political
                )
                return RiskGateResult(
                    block = RiskBlock(
                        available = true,
                        scores = scores,
                        transcript = if (verbose) {
                            ClassifierTranscript(model = model, input = text)
                        } else null
                    ),
                    gated = scores.overall > threshold
                )
            }
        }

        if (classifier == null) {
            return RiskGateResult(
                block = RiskBlock(
                    available = false,
                    reason = "no classifier endpoint configured",
                    transcript = if (verbose) ClassifierTranscript(model = model, input = text) else null
                ),
                gated = true
            )
        }

        return when (val result = classifier.classify(text, verbose)) {
            is ClassificationResult.Success -> {
                if (contentHash != null) {
                    repo.insertVerdict(
                        contentHash, model, result.scores,
                        Clock.System.now().toEpochMilliseconds()
                    )
                }
                RiskGateResult(
                    block = RiskBlock(
                        available = true,
                        scores = result.scores,
                        transcript = result.transcript
                    ),
                    gated = result.scores.overall > threshold
                )
            }
            is ClassificationResult.Unavailable -> RiskGateResult(
                block = RiskBlock(
                    available = false,
                    reason = result.reason,
                    transcript = result.transcript
                ),
                gated = true
            )
        }
    }
}

RiskClassifier — Ollama HTTP client

This users Ollama's /api/chat with format pinning the JSON schema and =options.temperature=0= for determinism. 45s request timeout: cold model loads on the 4090 are slow.

The Prompt

kotlin#+name: classifier-system-prompt
"You are a privacy risk classifier. " +
"Classify the following text against five categories, scoring each 0.0 to 1.0 " +
"for how much of that category's sensitive content the text contains. " +

"Categories: " +
"health (personal medical/health information), " +
"private (intimate or deeply personal non-medical information), " +
"financial (financial details, salaries, account information), " +
"employer (work-owned or employer-confidential data), " +
"political (politically sensitive opinions or affiliations). " +

"Respond with ONLY a JSON object with keys " +
"health, private, financial, employer, political."

The Client

kotlin#+name: classifier-client:noweb yes
class RiskClassifier(
    private val endpoint: String,
    private val model: String,
    private val httpClient: HttpClient,
    private val timeoutMs: Long = 45_000
) {
    private val categories = listOf("health", "private", "financial", "employer", "political")

    private fun formatSchema(): JsonObject = buildJsonObject {
        put("type", JsonPrimitive("object"))
        put("properties", buildJsonObject {
            categories.forEach { cat ->
                put(cat, buildJsonObject { put("type", JsonPrimitive("number")) })
            }
        })
        put("required", JsonArray(categories.map { JsonPrimitive(it) }))
    }

    private fun requestBody(text: String): JsonObject = buildJsonObject {
        put("model", JsonPrimitive(model))
        put("stream", JsonPrimitive(false))
        put("format", formatSchema())
        put("options", buildJsonObject { put("temperature", JsonPrimitive(0.0)) })
        put("messages", JsonArray(listOf(
            buildJsonObject {
                put("role", JsonPrimitive("system"))
                put("content", JsonPrimitive(
                    <<classifier-system-prompt>>
                ))
            },
            buildJsonObject {
                put("role", JsonPrimitive("user"))
                put("content", JsonPrimitive(text))
            }
        )))
    }

    /**
     ,* Classify [text]. Never throws: every failure mode returns
     ,* [ClassificationResult.Unavailable] so callers fail closed. When
     ,* [verbose] is set, the request/response exchange rides on the result
     ,* as a [ClassifierTranscript].
     ,*/
    suspend fun classify(text: String, verbose: Boolean = false): ClassificationResult =
        withContext(Dispatchers.IO) {
            val request = requestBody(text).toString()
            fun transcript(rawResponse: String? = null) = if (verbose) {
                ClassifierTranscript(
                    model = model,
                    input = request,
                    rawResponse = rawResponse
                )
            } else {
                null
            }
            try {
                val response = httpClient.post("$endpoint/api/chat") {
                    contentType(ContentType.Application.Json)
                    setBody(request)
                }
                if (!response.status.isSuccess()) {
                    return@withContext ClassificationResult.Unavailable(
                        "classifier http ${response.status.value}",
                        transcript("HTTP ${response.status.value}")
                    )
                }
                val raw = response.bodyAsText()
                val body = Json.parseToJsonElement(raw).jsonObject
                val content = body["message"]?.jsonObject?.get("content")?.jsonPrimitive?.contentOrNull
                    ?: return@withContext ClassificationResult.Unavailable(
                        "classifier response missing message.content",
                        transcript(raw)
                    )
                val scores = try {
                    ToolJson.decodeFromString(RiskScores.serializer(), content.trim())
                } catch (e: Exception) {
                    return@withContext ClassificationResult.Unavailable(
                        "classifier verdict not parseable",
                        transcript(content)
                    )
                }
                ClassificationResult.Success(scores.clamped(), transcript(content))
            } catch (e: Exception) {
                ClassificationResult.Unavailable(
                    "classifier unreachable: ${e.message}",
                    transcript()
                )
            }
        }
}

/** Coerce every category to [0, 1] and treat NaN as 0. */
fun RiskScores.clamped(): RiskScores = RiskScores(
    health = clamp01(health),
    privateScore = clamp01(privateScore),
    financial = clamp01(financial),
    employer = clamp01(employer),
    political = clamp01(political)
)

private fun clamp01(v: Double): Double = when {
    v.isNaN() || v < 0.0 -> 0.0
    v > 1.0 -> 1.0
    else -> v
}

RiskClassifier.kt Composition

kotlin#+name: classifier-composed:tangle ../src/jvmMain/kotlin/computer/whatthefuck/arcology/llm/RiskClassifier.kt:noweb yes
package computer.whatthefuck.arcology.llm

import computer.whatthefuck.arcology.utils.HashUtils
import io.ktor.client.*
import io.ktor.client.request.*
import io.ktor.client.statement.*
import io.ktor.http.*
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.withContext
import kotlinx.serialization.SerialName
import kotlinx.serialization.Serializable
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.JsonArray
import kotlinx.serialization.json.JsonPrimitive
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
import kotlinx.serialization.json.jsonObject
import kotlinx.serialization.json.jsonPrimitive
import kotlinx.serialization.json.contentOrNull
import kotlin.time.Clock

<<classifier-types>>

<<classifier-client>>

<<classifier-gate>>

Types

private is a Kotlin hard modifier, so the category name survives to JSON via @SerialName on a privateScore property. The cache column is private_score for the same reason.

Risk Types

kotlin#+name: classifier-types
@Serializable
data class RiskScores(
    val health: Double = 0.0,
    @SerialName("private") val privateScore: Double = 0.0,
    val financial: Double = 0.0,
    val employer: Double = 0.0,
    val political: Double = 0.0
) {
    val overall: Double get() = maxOf(health, privateScore, financial, employer, political)
}

/** Result of one classification attempt. */
sealed interface ClassificationResult {
    data class Success(
        val scores: RiskScores,
        /** Set only when the caller asked for the transcript. */
        val transcript: ClassifierTranscript? = null
    ) : ClassificationResult
    /** Endpoint unreachable, timed out, or returned unusable output. Always fails closed. */
    data class Unavailable(
        val reason: String,
        /** Partial transcript when available (e.g. HTTP status, raw garbage). */
        val transcript: ClassifierTranscript? = null
    ) : ClassificationResult
}

/**
 * The risk block attached to every tool envelope. [available] false means
 * the classifier could not be consulted (fail closed); [scores] carry the
 * five category scores; [redacted] true means content was withheld.
 * [transcript] — the classifier's input/output exchange — is only attached
 * when the tool was invoked with =--verbose=.
 */
@Serializable
data class RiskBlock(
    val available: Boolean = true,
    val scores: RiskScores? = null,
    val redacted: Boolean = false,
    val reason: String? = null,
    val transcript: ClassifierTranscript? = null
)

/**
 * The full classifier exchange: what was sent (system+user prompt with the
 * text under classification) and what the endpoint answered. Debug artifact
 * for =--verbose= runs; never emitted otherwise.
 */
@Serializable
data class ClassifierTranscript(
    val model: String,
    val input: String,
    val rawResponse: String? = null
)

Tests

A fake Ollama endpoint (embedded Ktor CIO server on an ephemeral port) drives these; the fail-closed cases point at a port with nothing listening.

kotlin#+name: classifier-tests:tangle ../src/jvmTest/kotlin/computer/whatthefuck/arcology/llm/RiskClassifierTest.kt
package computer.whatthefuck.arcology.llm

import computer.whatthefuck.arcology.db.ArcologyDatabase
import io.ktor.server.engine.*
import io.ktor.server.cio.*
import io.ktor.server.response.*
import io.ktor.server.routing.*
import io.ktor.http.*
import io.ktor.client.*
import io.ktor.client.plugins.*
import kotlinx.coroutines.test.runTest
import kotlinx.serialization.json.JsonPrimitive
import kotlinx.serialization.json.buildJsonObject
import kotlin.test.*

class RiskClassifierTest {

    private fun makeRepo(): ToolQueryRepositoryImpl =
        ToolQueryRepositoryImpl(computer.whatthefuck.arcology.database.DatabaseFactory.createInMemoryDatabase())

    private fun client(): HttpClient = HttpClient {
        install(HttpTimeout) { requestTimeoutMillis = 5_000 }
    }

    /** Fake Ollama /api/chat returning the given verdict JSON as message.content. */
    private fun ollamaServer(verdictJson: String): Pair<EmbeddedServer<*, *>, Int> {
        // content is a JSON *string* whose text is the verdict JSON — build
        // with kotlinx so the nested quoting/escaping is done correctly.
        val responsePayload = buildJsonObject {
            put("model", JsonPrimitive("test"))
            put("message", buildJsonObject {
                put("role", JsonPrimitive("assistant"))
                put("content", JsonPrimitive(verdictJson))
            })
            put("done", JsonPrimitive(true))
        }.toString()
        // Ktor 3 dropped resolvedConnectors(); reserve an explicit ephemeral
        // port ourselves so tests can derive the URL without the connector API.
        val port = java.net.ServerSocket(0).use { it.localPort }
        val server = embeddedServer(CIO, host = "127.0.0.1", port = port) {
            routing {
                post("/api/chat") {
                    call.respondText(responsePayload, ContentType.Application.Json)
                }
            }
        }
        return Pair(server, port)
    }

    @Test
    fun `parses a healthy verdict`() = runTest {
        val (server, port) = ollamaServer("""{"health":0.1,"private":0.0,"financial":0.2,"employer":0.9,"political":0.0}""")
        server.start(wait = false)
        val classifier = RiskClassifier("http://127.0.0.1:$port", "test", client())
        val result = classifier.classify("some text")
        assertTrue(result is ClassificationResult.Success)
        assertEquals(0.9, result.scores.employer)
        assertEquals(0.9, result.scores.overall)
        server.stop(1000, 2000)
    }

    @Test
    fun `clamps out of range scores`() = runTest {
        val (server, port) = ollamaServer("""{"health":5.0,"private":-1.0,"financial":0.0,"employer":0.0,"political":0.0}""")
        server.start(wait = false)
        val classifier = RiskClassifier("http://127.0.0.1:$port", "test", client())
        val result = classifier.classify("some text")
        assertTrue(result is ClassificationResult.Success)
        assertEquals(1.0, result.scores.health)
        assertEquals(0.0, result.scores.privateScore)
        server.stop(1000, 2000)
    }

    @Test
    fun `garbage verdict fails closed as unavailable`() = runTest {
        val (server, port) = ollamaServer("""not json at all""")
        server.start(wait = false)
        val classifier = RiskClassifier("http://127.0.0.1:$port", "test", client())
        val result = classifier.classify("some text")
        assertTrue(result is ClassificationResult.Unavailable)
        server.stop(1000, 2000)
    }

    @Test
    fun `unreachable endpoint fails closed`() = runTest {
        val classifier = RiskClassifier("http://127.0.0.1:1", "test", client())
        val result = classifier.classify("some text")
        assertTrue(result is ClassificationResult.Unavailable)
    }

    @Test
    fun `gate fails closed without classifier`() = runTest {
        val repo = makeRepo()
        val gate = RiskGate(repo, classifier = null, model = "m", trustLocal = false)
        val result = gate.gate("some private content", threshold = 0.7)
        assertTrue(result.gated)
        assertEquals(false, result.block?.available)
        assertNull(result.block?.scores)
    }

    @Test
    fun `gate trusts local when flag set`() = runTest {
        val repo = makeRepo()
        val gate = RiskGate(repo, classifier = null, model = "m", trustLocal = true)
        val result = gate.gate("anything", threshold = 0.7)
        assertFalse(result.gated)
        assertNull(result.block)
    }

    @Test
    fun `gate caches verdict per content hash and model`() = runTest {
        val (server, port) = ollamaServer("""{"health":0.0,"private":0.0,"financial":1.0,"employer":0.0,"political":0.0}""")
        server.start(wait = false)
        val repo = makeRepo()
        val gate = RiskGate(
            repo, RiskClassifier("http://127.0.0.1:$port", "testmodel", client()),
            model = "testmodel", trustLocal = false
        )

        val first = gate.gate("salary negotiation notes", threshold = 0.7)
        assertTrue(first.gated) // financial 1.0 > 0.7

        // Fresh gate, dead classifier: still correctly scored via llm_verdict_cache
        val second = RiskGate(
            repo, RiskClassifier("http://127.0.0.1:1", "testmodel", client()),
            model = "testmodel", trustLocal = false
        )
        val result = second.gate("salary negotiation notes", threshold = 0.7)
        assertTrue(result.block?.available == true)
        assertEquals(1.0, result.block?.scores?.financial)
        server.stop(1000, 2000)
    }

    @Test
    fun `verdict cache is model-scoped`() = runTest {
        val repo = makeRepo()
        repo.insertVerdict("hash", "model-a", RiskScores(financial = 0.5), 1L)
        assertNull(repo.selectVerdict("hash", "model-b"))
        assertEquals(0.5, repo.selectVerdict("hash", "model-a")?.financial)
    }

    @Test
    fun `verbose off produces no transcript`() = runTest {
        val (server, port) = ollamaServer("""{"health":0.0,"private":0.0,"financial":1.0,"employer":0.0,"political":0.0}""")
        server.start(wait = false)
        val classifier = RiskClassifier("http://127.0.0.1:$port", "test", client())
        val result = classifier.classify("some text", verbose = false)
        assertTrue(result is ClassificationResult.Success)
        assertNull(result.transcript)
        server.stop(1000, 2000)
    }

    @Test
    fun `verbose on attaches request and raw response`() = runTest {
        val (server, port) = ollamaServer("""{"health":0.0,"private":0.0,"financial":1.0,"employer":0.0,"political":0.0}""")
        server.start(wait = false)
        val classifier = RiskClassifier("http://127.0.0.1:$port", "test", client())
        val result = classifier.classify("journal entry with salary figures", verbose = true)
        assertTrue(result is ClassificationResult.Success)
        val t = result.transcript
        assertNotNull(t)
        assertEquals("test", t.model)
        assertTrue(t.input.contains("journal entry with salary figures"))
        assertTrue(t.input.contains("\"financial\"")) // schema rides in the request
        assertEquals("{\"health\":0.0,\"private\":0.0,\"financial\":1.0,\"employer\":0.0,\"political\":0.0}", t.rawResponse)
        server.stop(1000, 2000)
    }

    @Test
    fun `verbose on failure carries raw junk as transcript response`() = runTest {
        val (server, port) = ollamaServer("""not json at all""")
        server.start(wait = false)
        val classifier = RiskClassifier("http://127.0.0.1:$port", "test", client())
        val result = classifier.classify("some text", verbose = true)
        assertTrue(result is ClassificationResult.Unavailable)
        // The raw unverifiable junk is exactly what verbose exists to expose.
        assertEquals("not json at all", result.transcript?.rawResponse)
        assertTrue(result.transcript?.input?.contains("some text") == true)
        server.stop(1000, 2000)
    }

    @Test
    fun `gate verbose live call carries transcript`() = runTest {
        val (server, port) = ollamaServer("""{"health":1.0,"private":0.0,"financial":0.0,"employer":0.0,"political":0.0}""")
        server.start(wait = false)
        val repo = makeRepo()
        val gate = RiskGate(
            repo, RiskClassifier("http://127.0.0.1:$port", "testmodel", client()),
            model = "testmodel", trustLocal = false
        )
        val result = gate.gate("doctor notes", threshold = 0.7, verbose = true)
        assertTrue(result.gated) // health 1.0 > 0.7
        assertNotNull(result.block?.transcript)
        assertNotNull(result.block?.transcript?.rawResponse)
        server.stop(1000, 2000)
    }

    @Test
    fun `verbose bypasses the verdict cache`() = runTest {
        val (server, port) = ollamaServer("""{"health":1.0,"private":0.0,"financial":0.0,"employer":0.0,"political":0.0}""")
        server.start(wait = false)
        val repo = makeRepo()
        val live = RiskGate(
            repo, RiskClassifier("http://127.0.0.1:$port", "testmodel", client()),
            model = "testmodel", trustLocal = false
        )
        live.gate("doctor notes", threshold = 0.7) // populates cache

        // Verbose + dead classifier: cache is bypassed, so the call hits the
        // dead endpoint and fails closed as unavailable (NOT a cache hit).
        val bypassed = RiskGate(
            repo, RiskClassifier("http://127.0.0.1:1", "testmodel", client()),
            model = "testmodel", trustLocal = false
        )
        val result = bypassed.gate("doctor notes", threshold = 0.7, verbose = true)
        assertEquals(false, result.block?.available)
        assertNotNull(result.block?.transcript)

        // Non-verbose same gate: cache hit serves the live scores.
        val throughCache = bypassed.gate("doctor notes", threshold = 0.7, verbose = false)
        assertEquals(true, throughCache.block?.available)
        assertEquals(1.0, throughCache.block?.scores?.health)
        server.stop(1000, 2000)
    }

    @Test
    fun `gate verbose without classifier emits config-only transcript`() = runTest {
        val repo = makeRepo()
        val gate = RiskGate(repo, classifier = null, model = "m", trustLocal = false)
        val result = gate.gate("some private content", threshold = 0.7, verbose = true)
        // No classifier endpoint: transcript reflects configuration, not content.
        assertEquals(false, result.block?.available)
        assertNull(result.block?.scores)
        assertNotNull(result.block?.transcript)
        assertEquals("some private content", result.block?.transcript?.input)
        assertNull(result.block?.transcript?.rawResponse)
    }
}

Related Modules