Skip to content

About the author

This documentation, and the AI Agent Eval Harness it describes, are the work of Waldemar Szemat, AI Solutions Architect and Fractional Head of AI.

Most GenAI pilots never reach production. Taking them there and keeping them running at scale is my focus: measurement-first AI systems, evaluation harnesses, AI governance, and retrieval-augmented generation for regulated domains such as healthtech.

This project is a public, trilingual proof of work: a measurement-first, cite-or-refuse conversational health agent for medication adherence, paired with a CI-gated evaluation harness, built and evaluated on 100% synthetic data. It is a capability and readiness reference, not a medical device. Read the executive summary for the short version.

  • LLM evaluation and judge calibration (CI-gated quality and safety gates)
  • AI governance and readiness (HIPAA, EU AI Act, NIST AI RMF, ISO/IEC 42001, SOC 2, MITRE ATLAS)
  • Retrieval-augmented generation (hybrid retrieval, parent-document retrieval, citation enforcement)
  • Healthtech and other regulated-domain AI

Part of the portfolio of Waldemar Szemat · szemat.pro
GitHub · LinkedIn