Blog
Insights on AI, security, software architecture, and building what's next for ambitious businesses.
2240 articles · Page 1 of 112
The Quarterly Agent Review: Expand, Fix or Retire Each Workflow
A practical quarterly checklist for deciding whether to expand, fix, pause or retire an AI-agent workflow—using operational evidence, not model enthusiasm.
Build an Agent Operations Dashboard: Metrics, Events and SQL
A practical tutorial for building an agent operations dashboard, from event schema and SQL to accuracy checks, failure analysis and useful metrics.
What Good AI Delivery Evidence Looks Like in a Case Study
A credible AI case study shows the baseline, timeframe, sample, failures and client-approved outcomes—not just a compelling result or model claim.
How to Measure an AI Operations Pilot Across Multiple Sites
A practical, site-by-site method for measuring AI answer accuracy, booking completion, corrections and staff time before deciding whether to expand.
Publishing a Credible AI Benchmark: PADISO’s Proposed Reporting Standard
A credible AI benchmark is an auditable record of data, versions, settings, costs and failures—not a leaderboard without a reproducible protocol.
An AWS Agent Landing Zone: The First Architecture Decisions
A practical AWS landing-zone design for agent workloads, covering account boundaries, identity, network paths, logging and cost allocation.
A Production Readiness Review for Microsoft Foundry Agents
A practical release checklist for Microsoft Foundry agents, covering security, quality, operations and cost gates before production.
Browser Agent Launch Review: A Go/No-Go Worksheet
A practical browser-agent go/no-go worksheet: define the release slice, collect evidence, test failure paths, and record a defensible launch decision.
When Should You Build Your Own Agent Harness?
When Should You Build Your Own Agent Harness?. Practical examples, tradeoffs and implementation guidance for technology leaders.
What a CTO Should Ask Before Approving the Next Frontier Model
A CTO’s release checklist for testing a frontier model against real workloads, operational limits and reversible production gates before approving a change.
When an Agent Should Stop and Ask a Person
A practical list of nine boundaries that tell AI agents when to pause, ask a person, and avoid acting on uncertain or irreversible decisions.
Why Search Console Clicks and GA4 Visits Disagree
Search Console clicks and GA4 visits measure different things. Align properties, dates, filters, tags and consent before treating the gap as a tracking defect.
Who Owns an AI Agent After Launch?
An AI agent needs a named business owner and operational owners after launch. Use this practical framework to assign decisions, incidents and ongoing work.
Dayrun’s MCP Connector: Asking About Bookings, Staffing and Takings
Dayrun’s MCP Connector: Asking About Bookings, Staffing and Takings. Practical examples, tradeoffs and implementation guidance for technology leaders.
How Many Test Cases Are Enough for an AI Model Comparison?
There is no universal test-set size. Choose cases by decision risk, task coverage and the uncertainty your comparison can tolerate.
Foundry vs AgentCore: Compare the Same Workload and Control Requirements
A practical Foundry-versus-AgentCore comparison: test the same workload, identity boundaries, failure paths and operating responsibilities before choosing.
Astra vs Sonnet 5.5 vs Opus 5.5: How to Run Your Own Comparison
A reproducible protocol for comparing Astra, Sonnet 5.5 and Opus 5.5 using your own tasks, measured results and a versioned decision record.
Sonnet 5.5 Migration: Build a Regression Suite Before Switching
A practical Sonnet 5.5 migration checklist: baseline prompts, tools and real tasks, define release gates, and make a reversible decision from evidence.
Cloudflare Kitesurf and WebMCP: What Changes for Browser Agents?
Cloudflare Kitesurf’s WebMCP update points to a new browser-agent interface, but beta limits and page-level variability matter for enterprise plans.
Sonnet 5.5 vs Opus 5.5: Which Work Should Your Team Route to Each?
A practical Sonnet 5.5 vs Opus 5.5 routing framework: compare task quality, escalation risk, latency and total cost using your own work.