Skip to content

Engineering an Evidence-Backed Portfolio Agent

Posted on:October 11, 2026

My portfolio agent now has two audiences: recruiters asking about my experience, and engineers inspecting how the system works. The first needs concise, grounded answers. The second needs evidence about retrieval, tool execution, validation and the system’s limits.

You can try the agent or inspect its architecture and latest results.

Quotes, tools and observable execution

A role assessment extracts the requirements from a supplied description and matches them against published experience. Structured JSON alone is insufficient: every supporting quotation must actually appear in the cited source. The Worker checks the source ID, status, quote, length and structure, with one bounded correction attempt. A matching quote still needs semantic evaluation to establish that it supports the requirement.

For implementation questions, tools read selected evidence files. This site’s repository is private, so visitors need website-hosted source snapshots rather than inaccessible GitHub links. The published files are explicitly allowlisted; their URLs contain a hash of the evidence snapshot. Deployments preserve these assets, so old citations can still be inspected. Other public projects use immutable GitHub revisions. Unavailable files remain an operational limitation.

The engineering view shows retrieval sources, observable tool calls, actual latency, token usage, reported LLM cost and model/prompt/corpus/code versions. It shows execution data, not internal model reasoning. Browser messages contain user/assistant text; the server owns instructions and tool results. Durable storage provides a shared request budget across Worker instances and restarts.

What the first real-provider run found

The initial suite ran on 11 October 2026 against source revision 881b2ea18b708517a2637f808803158e64c14183, before the website-hosted citation fix and expanded skills profile. It exercised the real Worker, embeddings, retrieval, models and tools. Browser behavior was checked separately with end-to-end tests.

MeasurementObserved result
Primary scenarios18 attempted; 10 passed, 2 failed evidence checks, 6 stopped with provider HTTP 403
Completed primary scenarios12 produced complete execution traces
Median first token, completed primary scenarios1.35 seconds
Median agent completion, completed primary scenarios7.99 seconds
Median total runner time, completed primary scenarios12.16 seconds, including network, follow-ups and model grading
Provider-reported LLM cost, entire run$0.312501, including the evaluated baseline and semantic grading; infrastructure/embedding costs excluded

The archived measurement data records versions, per-scenario outcomes, timings and reported cost without publishing conversations.

The two evidence failures were the portfolio’s private repository link and an unavailable Ratchet README lookup. Most importantly, provider HTTP 403 responses prevented the final scenarios from completing. Those results are errors, not successful evaluations. The run does not diagnose the exact provider account condition or establish the quality of the unexecuted cases.

One contextual RAG follow-up passed while its latest-message-only retrieval baseline missed the expected evidence. The remaining baseline cases were interrupted. This is a useful failure example, not enough data to claim a general or statistically significant retrieval improvement.

The automated judge checked grounding, answering the question and honest handling of evidence. It has not been calibrated against human labels. A small curated suite and a model grader can detect regressions; they cannot prove universal accuracy or independently verify employment achievements.

Changes driven by the evidence

The citation failure led to publishing selected source evidence directly on the website. A live assessment also added requirements that were absent from the job description, so its instructions now constrain it to the stated requirements. Missing documentation must not be treated as proof that a technology was never used. Confirmed skills belong in the skills list; role-specific examples and metrics come from actual work history.

The engineering page publishes current-version aggregate results. Evaluations run on demand in GitHub Actions, with versioned scenarios, downloadable owner reports and optional model grading. Reports from an older deployment are kept distinct from results for the current revision.

The next evidence to consolidate is from my MLOps work at Secret Escapes: observability, production monitoring, scaling and the lifecycle from training through batch inference and real-time serving. The goal is to make that ownership concrete through systems, decisions and outcomes.

Comparing inexpensive assessment models

A later screening run tested the same two role descriptions on four models, with identical evidence, exact-quote validation and one bounded correction. It cost $0.083110 including the fixed automated grader. All eight assessments passed structure and quotation validation.

ModelValidated casesCorrectionsAssessment cost, two casesMean end-to-end time
Haiku 4.52/20$0.02062413.07 s
GPT-5.4 Mini2/20$0.01304910.51 s
Gemini 3.8 Flash2/20$0.0141829.83 s
Qwen3.7 Plus2/21$0.00729731.73 s

Assessment cost excludes grading; end-to-end time includes it. The archived comparison measurements record the exact versions and automated outcomes without publishing conversations. Gemini was the fastest observed option without retries and cost less than Haiku; GPT Mini was close and slightly cheaper. Qwen was cheapest but slower, with one correction. This supports a production candidate for this workflow, not a general reliability ranking from two cases.

The grader also exposed its own failure: it flagged unsupported claims that an assessment explicitly declined to make. That automated failure is preserved and annotated in the archive. Grading now requires exact affirmative claim excerpts, checks verdict consistency, and allows one bounded correction. An unvalidated grade is reported as a grading error, with its cost retained. Testing the grader is part of testing the AI system; a plausible explanation alone is insufficient.

A full run exposed infrastructure and grader failures

The next full production run tested all 18 scenarios at revision 4a92951 for a reported $0.365, including grading. Fifteen passed. All four Gemini role assessments completed with valid quotations without a correction. The training assessment recognised Feast from the published skills, but the automated grader incorrectly demanded a demonstrated project. That is a grader disagreement, not evidence of a missing capability. Rubric v3 clarifies that a technology in an explicit skills list establishes a self-reported skill; specific implementation claims still need specific evidence.

Nagare and Ratchet README inspection received HTTP 403 from GitHub. I added a fallback that reads only allowlisted files at reviewed public commit hashes from the raw host when the API returns 403 or 429. Pinned evidence is labelled and its citation identifies the actual revision; a short cache allows another attempt to resolve the latest revision. Other retrieval failures are still reported candidly.

The archived full-run failures remain visible. A targeted follow-up checks only the three affected scenarios on the newer deployment; its scorecard is not a claim that all 18 cases were repeated. That distinction makes it possible to investigate failures while keeping paid evaluation costs bounded.