Codex/GPT-5.5 AgentHub Demo: A Single-Run Front-End Review

A single AgentHub front-end demo, preserving the Chinese prompt and screenshots alongside the author’s recorded interactions, historical build figures, and unverified production limits.

Codex/GPT-5.5AgentHub DemoSingle-Run ObservationReact PrototypeMobile ReviewEvidence Limits

Evidence and Method · English Summary

This English summary retains the author’s recorded observations and screenshots. The original runnable demo, dependency lockfile, complete build log, and interaction traces are not available in this repository; this edit did not rerun that demo. Build sizes, timings, and code-structure details below are historical notes, not independently reproduced measurements. The nine-item checklist is a manual demo record, not an automated pass rate or production acceptance test.

Codex is the execution tool; GPT-5.5 is the model name in the author’s record. No session model snapshot, reasoning settings, or complete call log is preserved here to verify the actual routing. Agent, token, cost, risk, and check values in the interface are local mock data. The official model page documents naming and configuration, but cannot establish the snapshot used in this historical session.

The complete prompt is preserved in Chinese. This page summarizes the observations in English; the Chinese brief is not a 1,500-English-word prompt. No short-prompt comparison or repeated sampling was performed.

Read the original Chinese prompt · Official GPT-5.5 model documentation

What I wanted to test was not just whether it could draw a nice screen. I wanted to see whether Codex/GPT-5.5, used as an AI coding agent, could deliver an AgentOps workspace that runs, can be clicked through, and is worth a real product discussion when the request is close to real front-end work.

What This Demo Record Examines

What was delivered?Review the recorded navigation, board, charts, filtering, and drawers, separating visible controls from implemented operations.
What evidence is available?Screenshots and the Chinese brief remain available. Build figures and interaction results are historical notes; the demo project is not available here for a rerun.
What can this establish?A reference for prototype review, not a foundation-model benchmark, context-window test, production audit, or comparison with untested products.

1. Background

I did not give it a one-liner like "build a pretty dashboard." If your question is whether Codex can ship a front-end prototype, whether AI can really build a usable SaaS console, what an AI agent collaboration platform should look like, or whether GPT-5.5 can land a solid Vite + React + TypeScript page, the model name is only part of the story. The prompt sets explicit boundaries, although this run does not isolate its contribution from the model or tool configuration. In this run, I constrained product context, page structure, interactions, responsive behavior, code quality, and delivery requirements all at once.

2. What It Delivered

The final output was a Vite + React + TypeScript single-page app that opens directly into the AgentHub console itself. It did not detour into a marketing page. It placed the agent list, metric cards, kanban board, charts, risks, and timeline inside one working surface.

Desktop screenshot: left-side agent roster, top KPI cards, central kanban, plus charts and risks in one screen. The information density and section completeness are already close to a demo-ready SaaS console.
9Delivery checks
8Recorded demo checks met
1Partially passed (mobile)
561KBMain JS chunk
Top navigationProduct name, project selector, time range switcher, theme toggle, team summary, and avatar are all present.
Agent listAll 5 agents are there, each with status, task summary, cost, token usage, and a progress indicator.
Core metricsActive Agents, Completed Tasks, Open Risks, Failed Checks, and Estimated Cost all include trend context.
Project kanbanBacklog, In Progress, Review, and Done are complete; cards include assignee, priority, status, and check/file signals.
ChartsBoth Token and Cost Burn and Throughput and Checks are implemented with legend and meaningful data dimensions.
Risks and timelineRisk severity, ownership, descriptions, and collaboration events are all implemented.
Detail drawerClicking an agent or task opens a drawer with description, activity, files, blockers, and action buttons.
Responsive behaviorA horizontal-rail card is partly visible at the right edge; scroll and touch reachability remain unverified.
Build verificationnpm run build passes. Vite reports a main chunk around 561KB, fine for demos but still needs bundle work for production.

3. Strengths Breakdown

First: it got the product shape right

It did not interpret "AI collaboration console" as a page that talks about a product. It built an actual workspace. Agent roster on the left, work surface in the middle, then kanban, charts, risks, and timeline below. You can tell at a glance this is for tracking delivery, not a promo page. That call is harder than it looks. Many AI dashboards still spend the first screen on marketing copy or hide core functionality in secondary views. This one does not.

Second: the mock data is not random filler

Planner, Frontend Builder, Backend Integrator, Test Runner, and Code Reviewer map to believable tasks and issues. Backend is blocked by a gateway fixture, Test Runner finds failing specs, Reviewer tracks coverage and a large diff. That makes the page feel like a project in motion instead of a template collage. More importantly, risk alerts (test coverage dropped, API contract changed, large diff requires review) line up with current agent states. These details make this demo more coherent; no broader sample was collected to establish how common that is.

Third: interaction coverage is above average for a demo

Time range switching updates metrics and charts. Clicking an agent filters the board and opens details. Risks can be filtered by severity. Tasks can move across board states. For a page running entirely on local mock data, this is beyond a static screenshot. One detail worth calling out: time switching does not only change chart data, it also updates metric values and trend labels, which shows these states are actually wired together.

4. Issues Breakdown

The status dropdown feels out of place

The kanban status control uses a native <select>. It works, but inside this dark console it looks out of tune. Task cards are otherwise fairly detailed, with priority tags, assignee badges, and check counters, then this one control suddenly drops the polish back to browser default UI. This is a visual-consistency preference in this output. Native selects have useful platform behavior; choosing a custom control also requires keyboard and touch validation.

Charts are complete, but not fully unified with the UI system

Chart data, legend, and tooltip are all functional, but typography, spacing rhythm, and color pacing are slightly disconnected from the card system. Axis labels keep Recharts defaults while cards use Space Mono. The result is still readable, just not yet that "grown from one design system" feeling.

The mobile screenshot identifies a boundary to investigate

The author recorded capturing this Playwright screenshot at a Pixel 5 viewport (393×851). A card at the right edge of the agent rail is only partly visible, and the navigation occupies substantial space. A still image cannot distinguish an intentional scrolling cue from unreachable clipping; inspect scrollWidth, scrolling reachability, and focus behavior. The mobile check remains partial rather than being generalized to a defect on every phone.

Pixel 5 viewport (393px): a partly visible right-edge card and tall navigation. Interaction checks are needed to determine whether any content is unreachable.

5. One Key Test: How Did It React to an Extra Interaction Request?

I added a realistic interaction sequence: move Add saved card fixture coverage from Backlog to In Progress, switch to 7 Days, then filter to high risk only. This combination checks three things at once: whether task status really updates the board, whether time range really drives data, and whether risk filtering is real filtering rather than visual hiding.

The result: status updates were immediate, the card moved from Backlog to In Progress, time range switch replaced metric and chart data, and risk filtering reduced the list to high severity items. That tells us the page is not hardcoded visuals. Key interactions are wired through React state. From the code shape, these are managed with top-level useState and render-layer filtering, which is a reasonable tradeoff for a demo.

The same test also exposed the boundary clearly: changes live only in local memory, a refresh resets state, there is no undo flow, timeline entries are not written back from operations, and the status control itself is not fully productized yet. These are not defects in demo terms. They are the concrete engineering work between a prototype and a production surface.

6. Conclusion: When Is Codex a Good Fit for This Kind of Work?

This detailed Chinese brief produced a Vite + React + TypeScript prototype useful for discussion. The author recorded a successful build, a main JS chunk around 561KB, and linked board, time-range, and risk-filter interactions. That supports trying a similar prototyping workflow, without establishing delivery reliability from one run.

The value of this output is the conversion of an explicit brief into components, mock data, and an interactive surface. The prompt specified content, behavior, and boundaries, but no short-prompt control was run; more wording cannot be assumed to guarantee this quality.

For production use, validate mobile and keyboard flows, persistence, real APIs, error handling, and resource loading. These are concrete follow-up tasks, not a defensible “70% complete” measurement. Decide whether to continue the prototype against your own acceptance criteria.

Method and Result Questions

Does this establish Codex’s general front-end ability?

It records a usable React prototype produced under one explicit brief. There were no repeated runs, recorded seeds, or multiple project samples, so it does not establish a general success rate or a percentage of engineering work replaced.

What do eight met checks and one partial check mean?

They are the author’s nine-item display checklist, not a test-suite pass rate. Mobile was partial: the screenshot shows a partly visible card at the right edge of a horizontal agent rail, which still needs a scroll-reachability check. Visible drawer buttons do not establish real business integration.

What do screenshots and build figures establish?

Screenshots support discussion of visible layout and content. The 561KB main JS chunk and successful build are historical notes. The original runnable demo project, lockfile, and full build log are not in this repository, and this edit did not rerun performance or build-time checks.

How long must the prompt be?

The preserved prompt is a detailed Chinese-language brief, not 1,500 English words. It sets page, interaction, design, and delivery requirements. There was no comparison across prompt lengths, so this run cannot establish a word-count threshold or percentage contribution.

Did AgentHub run real agents?

No. Tasks, risks, cost, token counts, and check status are local mock data for the console layout and state interactions, not evidence of actual multi-agent execution, billing, or test results.

What needs validation before product use?

Check full keyboard and touch flows, narrow-screen scrolling, persistence, operation logs, real APIs, and error handling. Then use bundle analysis to assess resources. These are follow-up tasks, not checks claimed as already completed.