Blog
Insights on AI, security, software architecture, and building what's next for ambitious businesses.
17 articles in Tutorial · Page 1 of 1
Build an Agent Operations Dashboard: Metrics, Events and SQL
A practical tutorial for building an agent operations dashboard, from event schema and SQL to accuracy checks, failure analysis and useful metrics.
Tool Changes Can Break Agents: Add Contract Tests to CI
Build CI contract tests for agent tools that catch schema, permission and error-behavior regressions before a changed tool reaches production.
Tracing AWS Agents Across Models, Tools and Business Systems
A practical AWS tutorial for connecting agent traces to verified task outcomes, diagnosing failures, and building an operational telemetry record.
Tracing a Foundry Agent from User Request to Business Outcome
Tracing a Foundry Agent from User Request to Business Outcome. Practical examples, tradeoffs and implementation guidance for technology leaders.
Human Approval in an AI Harness: Pause, Review and Resume Safely
Build an approval gate that binds a reviewer’s decision to one exact action, expires safely, survives retries and verifies the result after execution.
Agent-Generated SQL: A Test Pack for Permissions and Wrong Answers
A practical test pack for agent-written SQL: catch cross-tenant leaks, row-limit failures and incorrect joins before results reach users.
Benchmark Agent Tail Latency: Trace One Workflow End to End
A reproducible tutorial for measuring agent p50 and p95 workflow latency, attributing time to stages, and separating fast failures from successful completion.
Browser Agent Regression Tests: A Website-Change Fixture Pack
Build a repeatable Playwright fixture pack for browser agents facing renamed controls, expired sessions and layout shifts—then verify the business state.
Design an Agent Tool Contract with JSON Schema and Failure Codes
Design an agent tool contract with JSON Schema, scoped permissions, approval-bound execution, stable failure codes and validation cases.
Prompt Caching: Test Warmth, Expiry and Real Savings
A practical tutorial for measuring prompt-cache warmth, expiry, accuracy, access isolation and real savings instead of assuming a longer TTL pays off.
Playwright and AI Browser Control: A Hybrid Workflow Walkthrough
Playwright and AI Browser Control: A Hybrid Workflow Walkthrough. Practical examples, tradeoffs and implementation guidance for technology leaders.
Context Engineering for Agents: A Keep, Summarise or Retrieve Experiment
Run the same agent tasks with keep, summarise and retrieve policies. Build a repeatable harness that measures quality, cost, latency and provenance.
Superset on Kubernetes: Upgrade and Recovery Drills That Matter
A practical Superset-on-Kubernetes walkthrough for rehearsing Helm upgrades, validating migrations, and recovering safely when a release goes wrong.
Cost per Successful Agent Task: A Worked Cost Model
A worked, reproducible model for calculating agent cost per successful task—including retries, tool calls, human review and failed outcomes.
Build an AI Harness Run Manifest: State, Tools and Stop Conditions
A hands-on tutorial for designing an AI harness run manifest, with explicit state transitions, approval-bound actions, stop conditions and failure handling.
Durable Execution for Agents That Run for Hours
A practical tutorial for designing long-running agents that checkpoint safely, resume after failure, and handle retries without repeating harmful side effects.
Retries Without Duplicate Actions: Idempotency for AI Agents
Retries Without Duplicate Actions: Idempotency for AI Agents. Practical examples, tradeoffs and implementation guidance for technology leaders.