Case study 2 · Reference architecture

Hybrid Streaming RAG

Retrieval-augmented generation on streaming data that takes personal data, auditability and partial failure seriously.

Reference implementation

The problem

Retrieval-augmented generation on streaming data must handle personal data, auditability and partial failures. Most demos ignore all three.

The approach

  • 01
    Personal data never leaves the ingress raw. Recognised personal data is replaced with versioned HMAC tokens before chunking, indexing, inference, events or audit writes.
  • 02
    Traceable requests. Every request gets server-issued request and correlation IDs and a W3C trace context. These travel in events and Kafka headers, and caller-supplied values are discarded.
  • 03
    Defined failure behaviour. If a downstream system fails, the workflow ends in a defined state, completed or compensated, with a matching compensating event.
  • 04
    Immutable audit through object-lock storage on S3, Azure or GCS.
  • 05
    Operational controls: bulkheads, rate limits, and Prometheus metrics with bounded cardinality.
  • 06
    Formal architecture documents: a TOGAF Architecture Definition and C4 models from Level 1 to Level 4.
flowchart LR
  C["Client"] --> A["API ingress<br/>request_id · correlation_id<br/>W3C traceparent"]
  A --> T["PII replaced with<br/>versioned HMAC tokens"]
  T --> K["Chunk and index"] --> V["Vector store adapter<br/>Chroma or OpenSearch"]
  T --> M["Inference / risk scorer<br/>adapter: Ollama"]
  A --> E["Kafka events<br/>with correlation headers"]
  A --> U["Immutable audit<br/>object lock: S3 · Azure · GCS"]
  A --> L["Lifecycle<br/>COMPLETED or COMPENSATED"]
Request path and failure handling, from the repository README. · Source: README

Honest scope

The README itself says this is a bounded compensating-transaction design, not full event sourcing.Repository →