StateSet iCommerce CUA Engine

If a personcan use it,your agent can too.

Automate work across APIs, browsers, desktops, and legacy portals with agents that can see, act, verify, recover, and ask for approval before consequential changes.

Agent runtime ready
95.8% accuracy on graded web tasksVerified actions with recoveryYour cloud or StateSet managed

Live automation run

RUN-1843 · Inventory reconciliation

Verified
admin.shopify.com/products/sku-1048

Detected discrepancy

SKU-1048 inventory mismatch

API value

0

Storefront

In stock

Correction approved

Set available inventory to 18

1

Inspect

Shopify inventory read

2

Compare

Storefront visual check

3

Approve

Inventory correction

4

Verify

Read-back matched

Use the best path

Call an API when one exists, run deterministic code when the rules are fixed, and use vision only when the work lives behind a screen.

MCP + APIs • code • computer use • model judgment

Verify the work

Check that every click, form submission, and data mutation produced the intended result before the agent marks the task complete.

Visual diff • read-back • retry • checkpoint

Govern every outcome

Put sensitive writes behind human approval and preserve screenshots, decisions, tool calls, cost, and results as run evidence.

Approvals • evidence • audit • permissions

Measured Capability

Performance you can inspect, not a demo claim.

The published graded web benchmark runs real browser tasks against known answers. Release governance turns benchmark and session evidence into a regression gate, then certifies the exact policy, runtime, audit state, and result.

95.8%

Graded web accuracy

23 of 24 independently graded tasks

$0.094

Average model cost

Measured per benchmark task

3

Execution engines

NSR, Claude, and OpenAI

4

Isolation tiers

Container, gVisor, microVM, and Kata

From Agent Demo to Operating System

Everything between a click and a trusted outcome

StateSet adds routing, verification, recovery, isolation, approvals, and release governance around computer-use models so they can run real operations.

See what operators see

Read dashboards, tables, fine print, dialogs, and legacy applications through screenshots and zoomed visual inspection.

The same agent can move from a structured API response to a portal screen without handing the workflow to another bot.

Act across every surface

Click, type, scroll, drag, use keyboard shortcuts, run code, edit files, and call typed MCP tools from one reasoning loop.

API-first execution keeps routine work fast; computer use closes the gaps where no reliable integration exists.

Recover instead of stalling

Visual verification, stuck detection, retries, circuit breakers, checkpoints, and self-healing keep long tasks moving.

The runtime distinguishes a completed action from a click that landed but changed nothing.

Scale in isolation

Run concurrent jobs in separate sandboxes with warm pools, tenant limits, scale-to-zero, and configurable isolation.

Each job gets its own desktop and execution context so parallel automation does not mean shared-session risk.

Keep humans in control

Require approval for consequential writes, show the proposed change, and verify the result after execution.

Tool guards, permissions, prompt-injection defenses, and fail-closed production checks enforce the operating boundary.

Deploy in your cloud

Run the platform in Azure, AWS, GCP, or your own Kubernetes environment with your identity, secrets, and data controls.

Terraform, Helm, keyless workload identity, pluggable storage, and multi-provider inference support enterprise deployment.

Execution Loop

Understand, route, act, verify, and govern

Every task follows the same accountable path, even when it crosses APIs, local code, browser tabs, and desktop applications.

Turn an operating request into a bounded objective.

1. Understand

The agent reads the task, gathers only the context it needs, identifies the systems involved, and defines what a successful outcome must look like.

  • Route commerce, support, onboarding, and general tasks to reusable skills.
  • Inspect current desktop and application state before acting.
  • Preserve attention with just-in-time retrieval and structured memory.

The task starts with an explicit objective and completion condition, not an open-ended instruction to click around.

Control flow
1
Natural-language task
2
Relevant context
3
Outcome contract
Commerce Workloads

Automate the work between your systems

The strongest computer-use workflows combine clean integrations with the screens, portals, and exception paths APIs cannot reach.

Catalog operations

Operational problem

Product data is spread across APIs, spreadsheets, admin screens, and storefront pages, so API-only checks miss what customers actually see.

How StateSet solves it

Audit catalog hygiene through APIs, use code for batch analysis, inspect storefront reality with computer use, and route proposed fixes through approval.

Outcome profile

  • Find corrupt SKUs, missing categories, junk tags, and inventory mismatches.
  • Normalize brand taxonomy and structured product attributes.
  • Reconcile system-of-record values against the live storefront.
Beyond Traditional Automation

Build around outcomes, not brittle click scripts

Traditional RPA works when the path never changes. StateSet combines adaptive reasoning with the operational controls required for dynamic interfaces and exception-heavy work.

How work is defined
Selectors, scripts, and rigid step-by-step bot flows.
Outcome contracts, reusable skills, and adaptive planning.
How systems are accessed
One integration or automation project per application.
API-first, deterministic when possible, GUI fallback when needed.
What happens after a UI change
The bot breaks until a maintainer rewrites selectors.
Visual reasoning re-evaluates the screen and verifies the new state.
How failures are handled
Retry the same step or send an opaque job failure.
Detect no-ops, re-plan, recover from checkpoints, or escalate with evidence.
What operators can audit
Logs of scripted steps and application errors.
Approvals, screenshots, decisions, tool timeline, cost, and verified result.
Start With One Outcome

Put one manual workflow on an evidence-backed automation path.

Choose a repetitive, measurable workflow. StateSet maps the API and GUI steps, defines the approval boundary, runs a graded pilot, and gives you the evidence to decide what should scale next.

Pilot success path

1

Define the outcome and a human-reviewed gold set

2

Map APIs, deterministic steps, and GUI-only gaps

3

Set approval, permission, cost, and isolation policy

4

Grade accuracy, evidence quality, and cost per outcome