IKA YOUR PRODUCT PARTNER
Back to portfolio ↗
06 / 09 · AI Quality & Trust Product

Proofline

The answer should be able to show its work.

When a support question touches payments, APIs, policies or account behavior, a fluent answer is not enough. Proofline retrieves approved evidence, shows what supports the answer, refuses when the evidence is insufficient, and hands a complete case to a human when the question is account-specific or high-risk.

AI QUALITY ENGINEERING

How should AI answer when sounding right is not good enough?

This case goes deeper into knowledge lifecycle, retrieval quality, refusals and how trustworthy support operates six months after launch.

01 / THE PRODUCT PROBLEM

Start with the decision the user cannot make well today.

When a support question touches payments, APIs, policies or account behavior, a fluent answer is not enough. Proofline retrieves approved evidence, shows what supports the answer, refuses when the evidence is insufficient, and hands a complete case to a human when the question is account-specific or high-risk.

The portfolio test: a client should understand the human problem before seeing a model, agent or framework.
02 / HOW RIKA THINKS

From ambiguity to a decision system.

The five-step reasoning pattern stays consistent; the actual product logic is specific to this problem.

FOCUS

What is the user buying?

Not chat. A trustworthy resolution.

FRAME

What must be true?

The answer must be applicable, current, cited and safe to act on.

EXPOSE

Where can trust break?

Bad source, weak retrieval, stale version, PII, prompt injection, overconfidence.

MAKE

What should AI own?

Interpret, retrieve, synthesize and explain; policy controls scope and escalation.

REFOCUS

What should improve?

Failures become corpus fixes, retrieval changes, prompt changes or new regression cases.

03 / TRANSFERABLE PROBLEMS

The pattern travels. The constraints change the solution.

This is how I would reuse the thinking without copying the product.

TRANSFERABLE PROBLEM

Healthcare information assistant

Same evidence/grounding pattern; clinical safety and regulation tighten the autonomy boundary.

TRANSFERABLE PROBLEM

Legal / policy copilot

Same citation and versioning problem; jurisdiction and precedent become critical metadata.

TRANSFERABLE PROBLEM

Internal enterprise search

Same retrieval problem; access control and confidential data dominate.

TRANSFERABLE PROBLEM

Developer copilot

Exact API versions, identifiers and code compatibility matter more than broad semantic similarity.

04 / USERS + BUYER

Who receives value, who operates it, and who pays?

AI products often fail when one persona is treated as ‘the user.’

Primary

Merchant / developer asking a product question

Wants a verifiable answer quickly, without needing to understand retrieval or model behavior.

Secondary

Support agent / support engineer

Needs the same evidence, classification and conversation summary when escalation occurs.

Buyer

Head of Support / Developer Experience

Wants safe deflection, lower handling effort and visibility into knowledge gaps without increasing wrong-answer risk.

JOBS TO BE DONE

Functional, emotional and social value.

Functional JTBD

Resolve product/policy/API questions from approved evidence and escalate the questions that need account state or human judgment.

Emotional JTBD

Reduce anxiety created by vague or contradictory support answers.

Social JTBD

Let the user confidently share or act on an answer because its source is visible.

CUSTOMER JOURNEY

The question changes as evidence arrives.

01

Ask

User states the question.

PII/scope gate runs first.

02

Retrieve

Find exact + semantic evidence.

Version/freshness filters remove invalid context.

03

Judge

Decide if the question is answerable.

Confidence is not only model probability; evidence quality matters.

04

Answer / refuse

Generate cited response or explain why the system cannot safely answer.

Account-specific questions prepare handoff.

05

Learn

Feedback + ticket outcome become knowledge/eval input.

Fix corpus, retrieval or policy—not blindly the prompt.

05 / PRODUCT DECISION RECORDS

The product is the sum of choices and trade-offs.

Each decision includes an implicit reversal test: better evidence can change the choice.

A refusal can be a successful outcome

If the evidence cannot support the claim, the product should preserve trust rather than maximize deflection.

Hybrid retrieval beats one search mode

Error codes and API identifiers need exact retrieval; natural-language questions benefit from semantic retrieval.

Knowledge operations is part of the product

Source ownership, freshness, deprecation and conflicting documents determine quality months after launch.

Account-specific answers require a different authority boundary

Public documentation can support policy-level answers; account data should require authenticated tools and stronger controls.

06 / AI SYSTEM

The architecture separates AI judgment from exact rules, evidence, tools and human authority.

Every connector has a defined job. Animated dashed lines represent active evaluation/learning loops rather than decorative motion.

Primary decision/data flowContext/evidenceHuman/consequential pathContinuous evaluation loop
Questionmerchant / developerScope + PIIsanitize · classify Lexical retrievalcodes · exact termsDense retrievalsemantic meaningMetadata filterversion · product · date Rerankerquery × evidenceAnswerabilitycoverage · conflict · freshness Grounded answeranswer + citationsHandoffsummary + evidenceEval + knowledge opsfailure → fix → reindex two retrieval lanescan this be answered safely?feedback / support outcomes → corpus · retrieval · policy · eval changes
HOW I WOULD BUILD IT

Point by point, from state to operation.

01

Ingest & version

Parse approved HTML/PDF/API schema; chunk at structure boundaries; attach URL, section, version and last-modified metadata.

02

Classify & sanitize

Detect scope, PII, exact identifiers and whether the question needs public knowledge or account state.

03

Retrieve twice

Lexical lane for codes/IDs + dense semantic lane for meaning; merge and metadata-filter.

04

Rerank

Cross-encoder reranks the short list so the model sees the most applicable evidence first.

05

Answerability judge

Check source coverage, freshness, contradictions and scope before generation.

06

Generate grounded answer

Produce short answer + steps + citations; only claims supported by context are allowed.

07

Verify & route

Faithfulness/citation checks + deterministic policy decide Answer / Clarify / Escalate / Refuse.

08

Operate the knowledge loop

Human corrections and unresolved query clusters trigger doc audits, new evals and re-indexing.

07 / EVALUATION ARCHITECTURE

How do I know the AI deserves to ship?

Component eval, system eval and product outcome are deliberately separated.

Component eval

Precision@K / recall@K

Precision@K / recall@K, reranker relevance, citation resolution, PII redaction, source freshness.

System eval

Faithfulness

Faithfulness, correct refusal, scope routing, handoff completeness and latency/cost.

Security eval

Prompt injection

Prompt injection, malicious retrieved text, unauthorized tool attempt and data leakage.

Product eval

Ticket deflection without repeat contact

Ticket deflection without repeat contact, satisfaction, escalation quality and unresolved knowledge-gap rate.

Release path

Offline eval → internal support alpha → sh

Offline eval → internal support alpha → shadow comparison → limited beta → monitored GA.

Regression loop

Thumbs-down/ticket outcome → failure taxon

Thumbs-down/ticket outcome → failure taxonomy → fix source/retrieval/prompt/policy → rerun suite.

Closed loop: trace → classify failure → add/refresh eval case → change source/retrieval/model/prompt/rule/tool → regression suite → controlled release → monitor outcomes/overrides → repeat.
08 / FEATURE PRD

A concrete example of how strategy becomes engineering work.

The PRD is intentionally feature-level and includes AI behavior, deterministic controls, telemetry and non-goals.

PRD example — Grounded settlement answer + safe escalation
Problem

A merchant asks why a settlement is delayed. Public policy can explain the general cycle, but the system must not invent account-specific status.

Outcome

Answer policy-level settlement questions from approved sources and identify when authenticated account data or human support is required.

User stories
  • As a merchant, I receive a short explanation with a source I can verify.
  • As a support agent, escalated cases arrive with intent, sanitized transcript, retrieved evidence and unresolved question.
Functional + AI requirements
  • Strip or mask sensitive fields before logging.
  • Classify public-policy vs account-specific question.
  • Run exact + semantic retrieval over approved/versioned corpus.
  • Rerank candidate evidence and reject stale versions.
  • Generate only from retrieved context and cite every factual claim.
  • If answerability is below policy threshold, clarify or escalate rather than fabricate.
Acceptance criteria
  • Every factual answer has at least one valid source.
  • No raw bank account/PAN/OTP/password is stored in logs.
  • Account-specific status is never invented from public docs.
  • Escalation contains top evidence and a three-bullet summary.
  • Prompt injection cannot override source/scope policy.
Telemetry

query_received · pii_redacted · retrieval_completed · citation_rendered · answer_refused · escalation_created · thumbs_down · ticket_within_24h

Non-goals
  • Account-specific transaction actions in V1
  • Using an LLM confidence number as the sole safety gate
  • Treating deflection alone as success
Engineering handoff

Typed state/schema · API/tool contracts · error states · permissions · eval fixtures · analytics events · rollout/rollback.

09 / BUILD VS BUY + RELIABILITY

Do not custom-build commodity infrastructure—and do not pretend every API always works.

BUILD DIFFERENTIATION

Trust policy, knowledge operations and evidence UX

Use standard infrastructure for storage, models, tracing and connectors where it does not create strategic advantage.

BAD-DAY MODE

Fallback to search + human support

Timeouts, stale data, partial results, rate limits and provider failures are explicit product states—not invisible backend details.

10 / WORKING PRODUCT JOURNEY

Use the product from input to changed state.

Every click represents a user/product decision and explains why the information is needed.

PROOFLINE · EVIDENCE-BACKED SUPPORT

Nothing has been answered yet.

Proofline does not begin by generating. It begins by deciding what evidence the question requires.

Public-policy scope detected.

LEXICAL
“settlement” / “delay”

Exact terms and policy headings.

SEMANTIC
Settlement timing

Meaning-based retrieval.

FILTER
Current version only

Old policy versions removed.

RERANK
Top evidence selected

Holiday exception ranks above generic cycle page.

Evidence is sufficient for a policy-level answer.

Settlements follow the documented cycle, with exceptions such as holidays or risk holds.

This prototype only demonstrates policy-level grounding. It does not claim to know this merchant's actual settlement status.

Grounded answer with visible source boundary.

ANSWER
Policy explained

Short answer first; uncertainty preserved.

CITATION
Settlement policy

Page + section + version metadata shown.

SCOPE
No account status invented

Public corpus cannot reveal private transaction state.

NEXT
Escalate if needed

Sanitized transcript + evidence package.

11 / PRODUCT LIFECYCLE

Autonomy and market scope are earned in stages.

STAGE 0

Corpus audit

Prove approved sources are complete enough.

STAGE 1

Internal copilot

Support team compares AI evidence to real resolutions.

STAGE 2

Customer self-service

Public-doc questions only; visible citations + escalation.

STAGE 3

Authenticated assist

Read-only account tools with stronger permissions and evals.

STAGE 4

Knowledge operations

Gap detection, doc-owner workflows and continuous quality management.

GTM + ADOPTION

A credible route into the market.

Who pays?

Support / developer-experience leadership replacing part of L1 handling cost.

Beachhead

High-volume documentation-answerable support categories.

Adoption

Embed before ticket creation but preserve easy human escape.

Expansion

Developer support → authenticated support → knowledge operations platform.

KILL / PIVOT CRITERION

What would make me stop?

Pivot or stop if the approved corpus cannot be kept current enough to support safe answers, or if users still require the same human investigation after seeing the grounded response.

12 / TOOL + MODEL DECISIONS

Tools are mapped to responsibility—not displayed as decoration.

Where the source material does not prove an implementation, the portfolio says proposed/candidate rather than “built with.”

OpenAI embeddingssource-backed PRD choice
Cohere reranksource-backed PRD choice
Claude Haiku / Sonnetsource-backed PRD choice
Supabase pgvectorsource-backed PRD choice
Freshdesk-style handoffsource-backed workflow
LangSmith / OpenTelemetryproposed observability pattern
13 / SECONDARY RESEARCH

Evidence establishes the problem environment. It does not magically validate the solution.

RESEARCH / EVIDENCE

NIST AI RMF + GenAI Profile

NIST frames trustworthiness and risk management across the AI lifecycle. Product implication: risk controls belong in design, evaluation and operations—not only at launch.

Open source ↗
RESEARCH / EVIDENCE

OWASP LLM01 Prompt Injection

OWASP notes that RAG does not eliminate prompt-injection risk. Product implication: retrieved content and user prompts need explicit trust boundaries and authorization controls.

Open source ↗
RESEARCH / EVIDENCE

RazorAssist authored PRD

The source PRD defines 500-token/50-overlap chunking, text-embedding-3-small, Cohere rerank-v3, Claude Haiku/Sonnet roles, pgvector, 0.7 routing and cited/refusal behavior. Proofline generalizes the reasoning without implying an official product.

Internal/source-project basis
EVIDENCE PACK

The working assumptions, PRD, eval suite and roadmap are inspectable.

No spreadsheet preview is embedded. Open the workbook only if you want the detail.

Download Proofline Product Evidence.xlsx ↗