Back to home

AI

How I connect AI business strategy and everyday use with responsible adoption and workload engineering.

How I approach AI

AI as an engineering discipline, not a demo: useful, governed, and measurable.

I apply the same discipline I use for cloud platforms to AI: start from a real problem engineers face, build on a governed foundation, and keep improving from measured outcomes. Day to day, that means building AI tools and agents that remove friction from engineering work, and making sure they run within clear identity, data, and cost boundaries.

CAF for AI (Microsoft Learn, opens in a new tab)

The Cloud Adoption Framework's AI scenario frames the organizational path: strategy, plan, ready, then govern, manage, and secure AI as an ongoing practice.

Well-Architected for AI (Microsoft Learn, opens in a new tab)

The Well-Architected AI guidance applies the five pillars to non-deterministic workloads, covering model training, hosting, inference, and evaluation.

The frameworks provide the questions; tool choices depend on the platform, data sensitivity, and what the team actually needs.

How organizations adopt AI

Based on the CAF AI adoption guidance (Microsoft Learn, opens in a new tab). The path applies whichever model provider or platform you choose.

Adopting AI isn't buying a model; it's a sequence of decisions about where AI creates value and how the organization keeps it under control. Like cloud adoption, the first steps run in sequence and the operational ones run continuously.

AI adoption · sequential

  1. 01AI strategy
  2. 02AI plan
  3. 03AI ready

Operations · in parallel

  • Govern AI
  • Manage AI
  • Secure AI
Adapted from the CAF AI scenario: build the foundation in order, then govern, manage, and secure AI for as long as it runs.
  1. AI strategy

    Which use cases are worth it? Pick problems with measurable value, choose between SaaS, PaaS, and IaaS AI, and set a data and responsible AI strategy.

  2. AI plan

    Does the organization have the skills, data, and access to deliver? Assess readiness, close skills gaps, and prioritize proofs of concept.

  3. AI ready

    Is there a foundation to build on? Establish AI landing zones with networking, identity, model access, and quota planned up front.

  4. Govern AI

    Are risks controlled? Define policy for model and data use, enforce it with guardrails, and track compliance and cost.

  5. Manage AI

    Is AI operated as a product? Standardize deployment, monitor quality and drift, and manage models and prompts through their lifecycle.

  6. Secure AI

    Is AI protected against new threats? Defend against prompt injection, data leakage, and model abuse with identity, network, and content controls.

Most AI risk is ordinary platform risk with new surfaces: reuse the cloud landing zone, identity, and cost controls rather than building a parallel stack for AI.

Responsible AI

Based on Microsoft's Responsible AI principles (Microsoft Learn, opens in a new tab). Principles only matter when they turn into checks in the delivery process.

I treat these six principles as review questions for every AI feature, answered before release and revisited as the system and its users change.

Fairness

Does the system treat similar people similarly? Evaluate outputs across user groups and fix disparities in data or prompts.

Reliability & Safety

Does it behave as intended under unexpected input? Test edge cases, filter harmful content, and keep a human in the loop for consequential actions.

Privacy & Security

Is data protected end to end? Minimize personal data, respect access boundaries in retrieval, and defend against prompt injection.

Inclusiveness

Can everyone use it? Design for different languages, abilities, and contexts, including bilingual teams like mine.

Transparency

Do users know they're working with AI and what it can't do? Disclose AI use, cite sources, and document limitations.

Accountability

Who owns the outcome? Assign owners, keep audit trails, and define how issues are reported and fixed.

In practice: write the answers down as part of the design review, automate what can be tested, and keep the rest as an explicit, owned checklist.

Designing AI workloads

Based on the Well-Architected AI workload guidance (Microsoft Learn, opens in a new tab). The same five pillars, applied to non-deterministic systems.

AI workloads behave non-deterministically and combine code and data into models, so quality has to be measured continuously rather than verified once. The pillar principles below adapt WAF to model training, hosting, and inference.

Reliability

5
  • Mitigate single points of failure

    Build redundancy into every critical component, preferring platforms with built-in fault tolerance and high availability.

  • Analyze failure modes; use proven patterns

    Isolate failures with bulkheads, and handle throttling and transient errors with retries and circuit breakers.

  • Balance reliability across dependencies

    Hold the inference endpoint, data stores, and service APIs to the same targets, using zones or multiple regions where needed.

  • Design for operational reliability

    Keep model responses fresh with timely, automated retraining; justify offline training through cost-benefit analysis.

  • Design for a reliable user experience

    Load-test under concurrency surges, set expectations for wait times in the UI, and use asynchronous patterns.

Security

6
  • Earn user trust

    Apply content safety across the lifecycle, strip unneeded personal data from stores, indexes, and caches, and moderate content in both directions.

  • Protect data at rest, in transit, and in use

    Encrypt every data store, use TLS on every hop, and treat the model itself as a high-value asset that can leak training data.

  • Invest in identity and access management

    Apply RBAC or ABAC on control and data planes, and segment identities so users only reach content they are authorized to see.

  • Segment the design

    Use private networking to image, data, and code repositories, and isolate inference node pools from other workloads.

  • Test security continuously

    Test for unethical behavior on every change, include inference endpoints in routine testing, and run red-team exercises.

  • Reduce the attack surface

    Require authentication on every inference endpoint, including system-to-system calls, and prefer constrained APIs.

Cost Optimization

4
  • Model the cost drivers

    Estimate data volume, query volume, required throughput, and hidden dependency costs such as indexing skillsets.

  • Pay for what you intend to use

    Match tiers to real usage, use spot capacity for interruptible training, reserve GPUs for AI work, and benchmark SKUs.

  • Use what you pay for

    Watch utilization, stop exploration and training compute when idle, favor write-once-read-many storage, and assign cost owners.

  • Optimize operational costs

    Automate retraining, accept slightly older data where accuracy allows, and prune unused features from feature stores.

Operational Excellence

6
  • Foster continuous experimentation

    Build on DevOps, DataOps, MLOps, and GenAIOps, and agree early on what acceptable model performance means.

  • Minimize operational burden

    Prefer managed platform services over self-hosting to simplify orchestration and day-2 operations.

  • Automate monitoring, alerting, and audit

    Define quality metrics with data scientists early, track experiments, and raise actionable alerts.

  • Detect and mitigate model decay

    Test for drift automatically, alert when responses deviate, and adapt data processing and training as use cases change.

  • Deploy safely

    Choose side-by-side or in-place updates, test against quality targets before release, and plan for emergencies.

  • Evaluate the experience in production

    Collect user feedback and, with consent, conversation logs to measure real-world quality.

Performance Efficiency

4
  • Establish performance benchmarks

    Baseline model quality and platform performance, then re-evaluate continuously rather than testing once.

  • Size resources to performance targets

    Load-test to choose platforms and SKUs: GPUs for training and fine-tuning, general-purpose compute for orchestration.

  • Collect metrics and find bottlenecks

    Track pipeline throughput, search latency and relevance, orchestrator time, and user engagement, including multimodal latency.

  • Improve from production signals

    Automate metric collection, alerting, and retraining to keep the model effective over time.

Design areas

Technical
Application design · Application platform · Training data · Grounding data · Data platform
Operational
Workload operations · MLOps and GenAIOps · Testing and evaluation · Responsible AI

Tradeoff: the strongest security controls limit how much encrypted data can be inspected or logged, and GPU-heavy performance must be weighed against cost; right-size continuously and record these decisions.