Trusted AI

Trusted AI

Home
Governance
Observability
Resources
Archive
About

Observability

The Feedback Gap
How in-product flags, annotation programs, and RLHF-style pipelines turn feedback into safer, smarter AI systems
May 27 • Jon Knisley
Evaluating Generative AI: Practical Tests for Safety, Bias and Fairness
How to assess toxicity, hallucinations, representational harms and group fairness
May 16 • Jon Knisley
Stop Guessing: How to Evaluate Modern LLM Applications
A practical framework for measuring what matters across prompts, retrieval, and tool behavior
May 11 • Jon Knisley
Your AI Passed Testing. Will It Survive Production?
Experiments, staged rollouts, and runtime guardrails that turn deployment risk into repeatable control
May 2 • Jon Knisley
Why Most LLM Evaluations Fail Before Production
How to build test sets, use synthetic data wisely, and trust LLM-as-a-judge without fooling yourself
Apr 25 • Jon Knisley
The Wrong Scorecard: Why Most Organizations Are Failing at AI Evaluation
A decision framework for investing evaluation resources where they matter most
Apr 19 • Jon Knisley
Your AI is Running. But is It Working?
The observability stack that makes AI systems debuggable, auditable, and improvable
Apr 9 • Jon Knisley
From ML Monitoring to AI Observability: What Changes with LLMs and Agents
Closing the visibility gap that separates AI leaders from liabilities
Mar 31 • Jon Knisley
© 2026 Jon Knisley · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture