IKA YOUR PRODUCT PARTNER
Back to portfolio ↗
03 / 05 · Travel disruption recovery

GuardianAI

A disruption changes more than one booking. It changes the dependency graph of the whole journey.

Travel data is fragmented across status, inventory, weather, maps, rail, hotel, visa/policy and payment systems. The environment can continue changing while the recovery is in progress.

Healthy tripbooked journey Disruptiondelay / cancel Recoveryfeasible plans Controlact / ask / escalate Recoveredkeep watching
01 / IN PLAIN ENGLISH

You should not need to be a product manager to understand this.

A delayed flight can break a connection, hotel arrival, airport transfer and budget at the same time. Guardian watches the trip, finds safe alternatives, asks before consequential changes, and keeps monitoring after the recovery.

What I want a client to understand: the product idea is only one part of the work. The value is in how the problem is framed, what evidence is trusted, where automation stops, and how the system learns after real outcomes.
POSITIONING

GuardianAI is not a list of replacement flights. It is a bounded-autonomy system that protects a changing trip while keeping consequential travel decisions under explicit control.

The point is to make the product's wedge obvious before the reader encounters any AI terminology.

PEOPLE + JOB TO BE DONE

Who experiences the problem — and what are they really hiring the product to do?

I separate the end user, secondary operator and buyer because the same product can create very different value for each.

PRIMARY USER

Traveler whose active trip is materially at risk

Needs a small number of viable choices, deadlines and control over costly/irreversible changes.

Why it matters: Under disruption, certainty and control matter more than browsing.

SECONDARY USER

Airline / OTA / TMC support or operations agent

Needs the same trip graph, impact analysis and prepared recovery options.

Why it matters: Agent assistance can reduce repetitive recovery work before customer autonomy expands.

BUYER / OWNER

Travel provider / TMC product & operations leader

Needs lower handling time and better recovery consistency without unauthorized transactions.

Why it matters: Early GTM is safer in shadow/agent-assist workflows.

JOBS TO BE DONE

Functional value is only part of the job.

The product also has to resolve an emotional tension and a social consequence.

Functional

Detect what a disruption breaks, find feasible recovery and safely complete or prepare the change.

This is part of the same job, not a separate “nice to have.”

Emotional

Reduce panic and restore control when the original trip no longer works.

This is part of the same job, not a separate “nice to have.”

Social

Help me communicate confidently with companions, employer or family that I have a credible recovery plan.

This is part of the same job, not a separate “nice to have.”

EMPATHY MAP

What is happening in the user's head and behavior?

Where the source workbook labels an item as a hypothesis, the portfolio keeps that label honest until real research replaces it.

THINKS

“What will fail next, and how much time do I have?”

FEELS

Urgency, loss of control and skepticism toward changing recommendations.

SAYS

“Do not tell me the flight is delayed; tell me if I will make the connection.”

DOES

Refreshes airline, weather, maps, hotel, rail, messages and support.

PAINS

Stale inventory, hidden policy, duplicate context, baggage/visa surprises.

GAINS

Early warning, viable options, preserved control, confirmed recovery.

CUSTOMER JOURNEY

The product follows the user's changing decision state.

Each stage exists because the user's question changes as new evidence enters the system.

01

Protect

Import itinerary + constraints.

Trip dependency graph starts monitoring.

02

Detect

Material disruption appears.

Impact translated into personal consequence.

03

Recover

Search + validate alternatives.

Hard constraints first, then ranking.

04

Approve & transact

Act / ask / escalate.

Booking uses idempotency + approval.

05

Verify & watch

Recovered trip is rechecked.

New disruption re-enters loop.

02 / HOW RIKA THINKS

From a messy problem to an inspectable decision.

This is the reasoning trail I would use as an independent product partner: focus on the user outcome, frame the system, expose the dangerous assumptions, make the smallest coherent product, then refocus using evidence.

FOCUS

What is the moment of value?

Before disruption becomes a missed trip, while viable choices still exist.

FRAME

What is the trip?

A dependency graph of bookings, time windows, policy and personal constraints.

EXPOSE

What makes autonomy dangerous?

Cost, irreversibility, stale inventory, visa/policy and low confidence.

MAKE

How should AI act?

Monitor, reason, use tools, compare recovery, then act only inside a permission envelope.

REFOCUS

How does it earn more agency?

Stable evals, low severe errors, successful rollback and subgroup quality.

03 / TRANSFERABLE THINKING

Similar problems show up elsewhere — but the difference matters.

A strong product partner does not copy a solution from one domain into another. I look for the shared problem pattern, then identify the constraint that changes the product decision.

ADJACENT PROBLEM

IT incident response

Similar: detect, diagnose, recover, verify. Different: the traveler is personally affected and transactions have cost/refund consequences.

ADJACENT PROBLEM

Healthcare scheduling disruption

Similar: cascading dependencies. Different: safety, clinical urgency and protected data dominate.

ADJACENT PROBLEM

Logistics re-routing

Similar: real-time alternatives and cost. Different: traveler preference, consent and human stress are first-class product inputs.

04 / PRODUCT DECISION RECORD

The product is shaped by the choices it refuses to hide.

These are not feature descriptions. They are decisions a client, engineer or operator can challenge.

0.7 is a prototype threshold, not universal truth

Confidence is one input to autonomy; it must be combined with cost, evidence and reversibility.

120% cost guardrail needs a baseline

The product compares against the cheapest viable recovery, not the original trip price.

Booking is not the end of recovery

Guardian revalidates the changed itinerary because another downstream commitment can still fail.

AUTONOMY POLICY

Model capability is not permission.

The 0.7 confidence and 120% cost rules are prototype routing assumptions used to demonstrate control. In production, they would be tuned against false-action cost, false-escalation cost, urgency and observed traveler outcomes.

Act

High confidence · low consequence · reversible · inside permission envelope.

Ask

Material cost, booking change or meaningful trade-off requires traveler approval.

Escalate

Low evidence, vulnerable traveler, visa/accessibility exception, conflict or irreversible consequence.

05 / HOW THE SYSTEM WORKS

The architecture explains who consumes what, where judgment sits, and how failure becomes learning.

Every connector has a defined job. Nothing is drawn simply to make the diagram look technical.

Primary decision/data flowRead-only evidence/contextHuman approval / consequential pathEvaluation / improvement loop
Trip graphbookings + dependenciesLive signalsstatus · weather · news Monitormaterial changeImpactbroken commitmentsRecoveryfeasible alternativesRisk + costconfidence · cost · reversibility Bookingprepare actionTraveler approvalconsequential actionVerify / watchre-check recovered tripEval looprouting · RAG · booking hard constraints + rankingapproval when consequence requirespost-action verificationoutcome / override → regression + threshold tuning
WHY THIS STRUCTURE

Separate uncertainty from authority.

AI can interpret and reason; deterministic services validate exact constraints; humans retain consequential judgment.

WHAT THE EVAL LOOP DOES

Failures become regression cases.

Traces, overrides and outcomes are classified so a retrieval/model/rule change can be tested before release.

WHAT A CLIENT CAN ASK

“Where can this go wrong?”

The answer should point to a node, a failure mode, an owner and a measurable control—not a generic “AI risk” statement.

Solid line

The normal forward path: a request, evidence packet, decision or approved action moves from one component to the next.

Static dashed line

A secondary validation or control path. It checks/qualifies the primary flow but is not continuously running.

Moving dashed line

An active monitoring/evaluation loop. Outcomes, overrides or changing state keep flowing back into checks, regression tests or the next decision.

HOW I BUILD THE AI SYSTEM

The agent is not the product. State, tools, policy and evaluation make the agent usable.

This is the implementation logic I would use with engineering: typed state first, clear contracts for each AI responsibility, deterministic rules for exact constraints, full tracing, and evidence-based expansion of autonomy.

01

Build trip graph

Bookings, connections, arrival deadlines, baggage, hotel and traveler constraints become a dependency graph.

02

Monitor signals

Operational, weather, news/geopolitical, historical and connection-geometry features update risk state.

03

Predict impact

Impact Agent maps an external event onto specific itinerary nodes and time-to-decision.

04

Search tools

Recovery Agent queries air/rail/hotel/ground inventory in parallel and keeps expiry/freshness.

05

Apply hard constraints

Visa, baggage, accessibility, minimum connection, budget and policy filters can reject plans before ranking.

06

Route autonomy

Prototype 0.7 confidence + 120% cost rule combine with reversibility and permissions: Act / Ask / Escalate.

07

Execute safely

Approved booking uses structured tool calls, idempotency and rollback/saga handling for partial failure.

08

Keep watching

Post-action verification re-reads bookings and feeds routing/RAG/transaction failures into evals.

HOW I WRITE THE EVALUATIONS

Evals are release criteria, not a scorecard added after the demo.

The dataset is designed around failure: happy paths, edge cases, missing data, conflicts, stale knowledge, adversarial inputs and high-consequence actions. Component quality, system quality and product outcome are measured separately.

1 · DATASET

Cancellation, severe weather, connection delay, stri

Cancellation, severe weather, connection delay, strike, geopolitical alert, rail disruption, inventory expiry, visa, baggage, news false-positive.

2 · DETECTION

Risk precision/recall, lead time and calibration; fa

Risk precision/recall, lead time and calibration; false alerts matter because they create anxiety.

3 · RECOVERY

Viable-plan recall, provider latency, policy freshne

Viable-plan recall, provider latency, policy freshness and hard-constraint rejection.

4 · ROUTING

Missed high-risk escalation is a critical error; ove

Missed high-risk escalation is a critical error; over-escalation is measured separately.

5 · TRANSACTION

Duplicate/unauthorized action must be zero; expiry/p

Duplicate/unauthorized action must be zero; expiry/partial-commit cases test rollback.

6 · CLOSED LOOP

Safe Recovery Rate + time-to-safe-plan + replan prec

Safe Recovery Rate + time-to-safe-plan + replan precision decide whether autonomy can expand.

The closed loop: trace → classify failure → add/refresh eval case → change prompt/retrieval/model/rule/tool → run regression suite → controlled release → monitor outcomes/overrides → repeat.
FEATURE PRD EXAMPLE

This is how I turn product judgment into something a squad can build.

The PRD carries the user problem, product outcome, AI/system behavior, data/API dependency, acceptance criteria, telemetry, safety boundary and non-goals.

PRD example — Disruption recovery approval
Problem

A 96-minute delay makes a connection infeasible and creates cascading hotel/ground constraints.

Outcome / objective

Surface the personal impact early, generate viable recovery plans, explain trade-offs and complete only the approved safe action.

User stories
  • As a traveler, I understand which parts of my trip are broken and how urgent the decision is.
  • As a traveler, I can approve a viable plan knowing cost, arrival impact, baggage and downstream feasibility.
Functional + AI requirements
  • Trip graph must identify all downstream commitments affected by the disruption.
  • Recovery search returns at least two viable plans when inventory permits.
  • Inventory/fare is revalidated immediately before transaction.
  • Risk router combines confidence, cost ratio, irreversibility, data freshness and user permission.
  • Plans above 120% of cheapest viable baseline require approval in the prototype policy.
  • Booking confirmation triggers post-action itinerary verification.
Acceptance criteria / Definition of Done
  • No plan violating visa/baggage/accessibility/minimum connection can be recommended.
  • Expired inventory cannot be transacted.
  • Every consequential booking change has explicit approval or signed permission.
  • Recovered itinerary remains monitored until stable.
Telemetry

material_alert · impact_opened · option_viewed · approval_requested · booking_confirmed · replan_triggered · human_override

Non-goals
  • Claiming 0.7/120% are production-optimized thresholds
  • Autonomous high-cost irreversible booking changes in early rollout
  • Ending the flow at ‘booking request sent’
Engineering handoff

Typed state/schema, API/tool contracts, error states, permissions, eval fixtures, analytics events, design states and rollout/rollback plan accompany the PRD.

06 / WORKING CUSTOMER JOURNEY

Click because you are making a product decision — not because the page needs another button.

Each step tells the reader why the input is required, what happens next, and what changes in system state.

GUARDIANAI · WORKING DISRUPTION JOURNEY

Your itinerary is healthy.

BLR → DOH · on time
healthy
DOH → CDG · 75 min connection
feasible
Paris hotel check-in
protected

The delay breaks the connection.

monitorimpactretrieve inventoryrecoverjudge
CONNECTION
21 min projected

Infeasible.

HOTEL
Valid until 23:30

Still recoverable.

COST BASELINE
€184

Cheapest viable recovery.

BAGGAGE
Through-ticketed

Constraint retained.

Two plans are feasible, but they are not equally safe.

PLAN A
Later DOH flight

€184 · confidence .84 · arrive +3h05.

PLAN B
Reroute via FRA

€251 · 136% of baseline · faster arrival.

Why Plan B cannot auto-run
It exceeds the 120% prototype cost guardrail. High cost + booking change means explicit traveler approval.

Guardian asks before changing a confirmed booking.

Rebook to QR-039 at 18:20?

€184 · hotel remains valid · baggage transfer supported · arrival 22:05.

The booking is confirmed. Guardian keeps watching.

transaction confirmedbooking ref updatedhotel recheckedmonitor recovered trip
Why the product does not stop here
Inventory, status or hotel feasibility can change again. Recovery is complete only while the new trip graph remains healthy.
FAILURE LAB

The strongest proof is often how the product behaves when things go wrong.

These cases are intentionally designed to expose the point where the system should clarify, refuse, route or re-evaluate rather than continue confidently.

Stale availability

A plan can be feasible when generated and invalid seconds later. Booking inventory must be rechecked immediately before action.

High-confidence, high-consequence action

Confidence alone does not grant authority; cost, reversibility and traveler vulnerability still matter.

Recovery cascade

A successful flight rebooking can break hotel, transfer, baggage or visa assumptions, so post-action monitoring stays active.

PRODUCT STRATEGY · GTM · ADOPTION

A useful product still needs a believable path into the market.

This is a proposed go-to-market hypothesis, not a claimed executed launch. The wedge is chosen by problem intensity, integration readiness, measurable value and the cost of being wrong.

BEACHHEAD

Airline/OTA/TMC operations and support where disruption recovery volume is high.

Agent-assist lowers risk and gives labeled decisions for evals.

PILOT

Shadow monitor + recovery recommendations on selected routes/providers.

Measure feasible-plan recall, handling time and false-alert cost.

ADOPTION

Expose proactive traveler alerts only after ops quality is stable.

Trust grows through clear options, freshness and explicit approval.

EXPANSION

Hold/price → assisted transaction → bounded reversible autonomy.

Commercial proof: lower handling time, recovery consistency and traveler effort.

TOOL / PLATFORM MAP

The stack is shown by responsibility, not as a logo wall.

Tiles labelled proposed/candidate are architecture choices, not claims that the product is already deployed with that tool.

LA
LangGraphproposed stateful orchestration
PR
Provider APIs / NDCshopping & orders
N8
n8nworkflow candidate
PO
PostgreSQLtrip state + audit
RA
RAG layerpolicy / visa / provider rules
OB
Observabilitytraces + evals + rollback
SECONDARY RESEARCH

External evidence should change a product decision — not decorate the case study.

These sources are used to validate the problem context, integration assumptions or competitive boundary. None of them are treated as proof that the proposed product itself works.

VERIFIED EXTERNAL SOURCE

McKinsey — Remapping travel with agentic AI

McKinsey explicitly discusses agentic AI helping rebook travelers during disruptions and freeing frontline staff from repetitive recovery work. This supports the agentic travel-recovery opportunity.

Open source ↗
VERIFIED EXTERNAL SOURCE

McKinsey — How agentic AI could transform travel

McKinsey describes agents that can plan end-to-end journeys and rebook disrupted flights, reinforcing the value of tool-using systems that can continue working as conditions change.

Open source ↗
VERIFIED EXTERNAL SOURCE

IATA NDC — Offers & Orders

IATA’s NDC standard provides an official airline retailing data-exchange model based on Offers and Orders. It supports the integration architecture for shopping/rebooking; it does not validate GuardianAI’s autonomy policy.

Open source ↗
07 / EVALUATION ANALYSIS

The chart answers a product question.

The workbook figures below are technical/synthetic evaluation—not production outcome. The purpose is to expose where system-level quality can break even when individual components look strong.

Technical evaluation view

95.0%
95.0%
95.0%
77.6%

What I would do with this

The technical scenario suite shows strong routing and booking confirmation while average confidence is materially lower. That supports a product design with explicit approval thresholds rather than treating successful tool calls as proof that the system should act autonomously.

Decision: do not release based only on the prettiest component metric. The end-to-end user outcome and the highest-consequence failure slice remain release gates.
WORKBOOK EVIDENCE

The underlying Excel evidence is attached and inspectable.

The page only uses metrics that the workbook actually supports. Synthetic research stays synthetic; technical scenario evaluation stays technical; neither is presented as production adoption.

Evidence rule: workbook numbers are prototype technical evaluation unless the workbook explicitly supports another evidence class. Real user validation remains a separate research step.
EVIDENCE STATUS

What is proven, what is prototype-tested, and what still needs reality.

Designed sophistication is not presented as production evidence. The next validation step is visible so a client can judge the maturity of the work honestly.

SOURCE-BACKED

Problem and system logic

Existing project material, workbooks and technical scenario suites support the current product/system framing.

PROTOTYPE-TESTED

Interaction and decision logic

The local prototype demonstrates the intended journey and control boundaries; it is not a live production integration.

TO VALIDATE

Real-world outcome

The scenario suite is technical prototype evaluation. The thresholds are not production-optimized. Real disruption cases and traveler comprehension tests would determine the eventual autonomy policy.

Where I would apply this thinking

If an AI system can take actions, autonomy should be designed from consequence, reversibility, evidence and permission—not simply from how capable the model looks.

What could change my mind?

Real workflow observation, user behavior, production telemetry, economic evidence or a simpler alternative that achieves the same outcome with lower risk. The decision record should be reversible when better evidence appears.