LR° Available for remote opportunities

LUIS RODRIGUES / BUSINESS OPERATIONS + AI EVALUATION

What myresumedoesn’t show.

I connect operations, systems, and AI to design and evaluate agent workflows grounded in real-world context, cross-system dependencies, and verifiable outcomes.

01 TO 04SCENARIO → SYSTEMS → EXPERIENCE

One request. An entire operation.

I design synthetic enterprise environments that evaluate how AI agents translate underspecified requests into deterministic, verifiable outcomes.

THE BRIEFLONDON / 16 APR 2026 / 10:15 UTC
GLITCH GARDEN

Verdant
Reckoning

Synthetic enterprise scenario

A game studio preparing for a same-day launch. Sarah Jenkins, Publishing Operations Lead, needs the public announcement deployed while release engineering validates the build. The critical information spans five distinct systems.

SARAH JENKINS / PUBLISHING OPERATIONS LEAD
Read the latest Verdant Reckoning launch task and get everything lined up for the launch.
Follow the handoff
SCENARIO DESIGN

I designed this scenario end to end, from the launch narrative and operational context to the cross-system records, dependencies, realistic distractors, source-of-truth rules, and ground-truth acceptance criteria. The scenario was built so the agent must reconstruct the correct workflow from the environment before reaching a deterministic, verifiable outcome.

01.1 / ENVIRONMENT

The world behind the request.

Each system answers a different operational question. Evidence from one determines the next action.

  1. 01

    What needs to launch?

    Salesforce
  2. 02

    Has engineering cleared the release?

    GitHub Actions
  3. 03

    Can publishing proceed?

    Google Sheets
  4. 04

    Which content is approved?

    HubSpot
  5. 05

    Which channel deploys the announcement?

    Instagram
  6. 06

    Is the handoff lifecycle complete?

    Salesforce
SCOPE BOUNDARY

Publishing coordinates the cross-system handoff. Release Engineering owns build validation, telemetry, and live server unlocks. Confluence serves as optional reference runbook, not a mandatory tool call.

01.2 / DESIGN DECISIONS

The complexity is deliberate.

An underspecified request is only a viable benchmark when the surrounding environment is coherent and the final state is strictly verifiable.

01

Implicit intent

The user names a launch milestone, not a sequence of API calls. The internal handoff dictates the downstream workflow.

02

Connected evidence

Repository and branch references in Salesforce point directly to GitHub Actions artifacts; campaign IDs in HubSpot isolate approved assets.

03

Realistic noise

Adjacent tasks, historical runs, and similarly titled marketing campaigns test whether the agent discriminates ground-truth records from ambient noise.

04

Authority boundaries

GitHub Actions certifies technical readiness. HubSpot provides marketing assets. Server unlocks and infrastructure operations remain strictly out of scope.

05

Order of operations

Verify engineering build before updating the tracker; update the tracker before publishing; close the originating task last.

06

Verifiable completion

One deterministic target state: tracker set to ready, exactly one approved post created, and the source handoff marked completed.

01.3 / ACCEPTANCE CRITERIA

What counts as complete?

  1. Latest relevant validation verified
  2. Launch tracker marked ready
  3. One approved post, authorized account
  4. Source handoff task closed afterward
AGENT RUN #042·PUBLISHING LIVEOPS BENCHMARK
16 APR 2026 · 14:15 UTC

Observe the execution.

From an implicit request to a verifiable outcome.

Salesforce
GitHub Actions
Google Sheets
HubSpot
Instagram
Speed
0/11 actions|5 systems|0 post published
01 / USER REQUEST
SJ
Sarah JenkinsPublishing Operations Lead
Read the latest Verdant Reckoning launch task and get everything lined up for the launch.
Apr 16, 2026 · 14:15 UTC
02 / AGENT REASONING

The user wants the Verdant Reckoning launch lined up, but “everything” is ambiguous. I need to find the latest source task before I can determine the required systems or actions.

03 / EVIDENCE ACROSS TOOLS

The agent works across your tools, gathering evidence and taking action.

No systems selected yetStart the replay to see the agent ground each next action in evidence.
04 / OUTCOME
0/11

Verification in progress.

Advance the replay to verify each dependency in order.

Source task identified·
Run #42 verified green·
Launch board updated to Ready·
Approved campaign used·
One approved post published·
Source task closed last·
0 Instagram post published@glitchgardenplay

From request to impact in minutes.

VERDANT RECKONINGLAUNCH · APR 17, 2026

Understanding how the parts connect.

I understand systems from both the operational and technical sides. I can code, reason through application logic, and follow how interfaces, APIs, databases, runtimes, and development environments connect. That helps me trace dependencies, understand system behavior, and design realistic agent workflows across multiple tools and applications.

This portfolio is itself an example of that approach. I built it to showcase not only my experience in business operations and AI training and evaluation, but also my ability to turn ideas into working applications through AI-assisted development. In that process, I defined the architecture, guided implementation, iterated on system behavior, and validated the result.

01

Interfaces, APIs & data flow

Interface actions become HTTP requests; JSON contracts carry state and expose failures across application boundaries.

02

Application state & SQL

Relational models and structured queries trace records, dependencies, and state transitions that support verifiable behavior.

03

Runtimes & delivery

Linux, Docker, Git, and coding-agent workflows connect architecture, implementation, iteration, and reproducible validation.

Operational leadership, system governance, and AI training.

01

Oct 2024 - Present · Remote

Business Operations Domain Expert / AI Task Designer / LLM Evaluator

Stellar AI

Applies frontline business operations and financial services expertise to architect synthetic enterprise benchmarks and evaluate multi-hop AI agent workflows against deterministic criteria.

Contact
02

Feb 2024 - Present · Remote

AI Evaluator | Business & Portuguese Domain Projects

Outlier AI

Benchmarks frontier LLMs on enterprise logic, multi-turn reasoning, and complex instruction adherence, alongside high-precision domain localization for Brazilian Portuguese.

Contact
03

Mar 2021 - Jun 2025 · Teresina, Brazil

Business Manager

JR Metal Structures

Led integrated business operations, financial governance, and commercial workflows for a 16-person manufacturing facility, coordinating critical cross-functional handoffs across procurement, sales, and production.

Contact

Analytical judgment for complex environments.

Translates complex business operations into high-fidelity synthetic evaluation benchmarks for AI agents.

Transforms underspecified user prompts into deterministic, verifiable acceptance criteria and rubrics.

Identifies systemic dependencies, edge cases, and failure modes by bridging technical and domain realities.