Claude Code AgentHub Demo: A Comparison With One Codex Run

Two AgentHub demo outputs from the same Chinese brief, with recorded interaction observations, mobile screenshots, historical build figures, and explicit limits on model attribution and comparability.

Claude CodeAgentHub DemoSingle-Run ObservationReact PrototypeMobile ReviewEvidence Limits

Evidence and Method · English Summary

This English summary retains the author’s recorded observations and screenshots. The original runnable demo, dependency lockfile, complete build log, and interaction traces are not available in this repository; this edit did not rerun that demo. Build sizes, timings, and code-structure details below are historical notes, not independently reproduced measurements. The nine-item checklist is a manual demo record, not an automated pass rate or production acceptance test.

Claude Code is the execution tool; its foundation model requires separate verification. The old “Claude Code 4.7” wording and the 4.7 URL segment are historical labels, with no CLI version output, session model ID, or request log here to establish the exact version. The opus-4.7, sonnet-4.6, and haiku-4.5 names in mock cards do not establish actual model calls. This review makes no context-window, token-usage, or model-tier capability claim from them.

The author recorded using the same Chinese brief for both runs, but full sessions are unavailable for a byte-for-byte comparison. Tool configuration, dependencies, and execution were not fully controlled. This compares two outputs without attributing their differences to foundation models. The Chinese section contains an excerpt; the complete Chinese prompt remains in the earlier review.

Read the original Chinese prompt · Official Claude Code model configuration

This is a second, independent demonstration of the same product brief. The earlier Codex/GPT-5.5 AgentHub review records the first output; this article preserves the Claude Code output and compares layout and interaction choices between those two prototypes.

What This Demo Record Examines

What was delivered?Review the recorded navigation, board, charts, filtering, and drawers, separating visible controls from implemented operations.
What evidence is available?Screenshots and the Chinese brief remain available. Build figures and interaction results are historical notes; the demo project is not available here for a rerun.
What can this establish?A reference for prototype review, not a foundation-model benchmark, context-window test, production audit, or comparison with untested products.

1. Experiment Setup

The author recorded starting from an empty directory and asking Claude Code to create a Vite + React + TypeScript project with the same Chinese product brief. The Chinese version includes a prompt excerpt, and the previous review preserves the full wording. Without full sessions and configuration records, this remains an output comparison rather than a controlled model experiment.

2. What It Delivered

Starting from an empty directory, Claude Code generated a Vite + React + TypeScript + Tailwind project, used lucide-react for icons and recharts for charts, without shadcn in the author’s record. The result was 9 feature components organized across 8 folders, plus a 15.3KB structured mock-data file.

Desktop screenshot: a 6-agent left roster, 5 KPI cards on top, stacked area + stacked bar charts in the middle, risks panel on the right, kanban and timeline below. Restrained dark palette with a teal accent — none of the "single blue-purple gradient" the prompt explicitly forbade.
9Delivery checks
9Recorded demo checks met
599KBMain JS chunk
2.42sProduction build
Top navigationAgentHub branding, custom project dropdown (3 projects with repo and branch), Today / 7 Days / 30 Days switcher, search button, theme toggle, active-agent counter, CL avatar.
Agent list6 agents (one extra Refactor Surgeon). Each carries a personal name (Atlas / Pixel / Forge / Sentry / Lens / Mason), role, status, mock model labels (opus-4.7 / sonnet-4.6 / haiku-4.5), current task, tokens, cost, and progress bar.
Core metricsAll 5 metrics implemented. Each card has a trend chip (up / down / warn / flat), delta label, and a contextual hint such as "Mason idle since 38m".
Project kanbanBacklog / In Progress / Review / Done columns. Cards include T-2041-style IDs, p0/p1/p2 priority pills, assignee avatar initial, file count, checks (passed/failed/pending icon counters), and relative time.
ChartsStacked area "Agent token usage" with a Tokens/Cost toggle; stacked bar "Task throughput" by completed/reviewed/failed. Both have grid, custom tooltip, and legend, fully aligned with the card design system.
Risks and timeline6 risks with multi-select severity pills (high/medium/low), hover-revealed dismiss button. Timeline includes a connecting vertical line and per-status icons.
Detail drawerTwo drawer modes for agent vs task: status pill, current-task progress, three-stat row, recent activity, open files, active risks, and a visible Pause / Reassign / View runs / Restart row; Pause/Reassign have no business handlers. ESC closes; opening locks body scroll.
Responsive behaviorSidebar + main on desktop; mobile flips to a top agent strip + single-column stack, kanban becomes 1 column, 2 at md, 4 at xl. The author reported no whole-page overflow at 393px; horizontal-rail reachability still needs an interaction check.
Build verificationnpm run build finished in 2.42s — 2,388 modules, main JS 599.24KB (gzip 170.40KB), CSS 26.49KB. Vite warned that the chunk exceeds 500KB, same caveat as the GPT-5.5 run.

3. What It Got Right

First: different controls and a different narrow-screen layout

The two observations discussed in the GPT-5.5 review were a native <select> for status changes and a partly visible right-edge card whose scrolling reachability remained unverified. Under the same prompt, Claude Code spontaneously built a custom dropdown menu for status moves (with a "Move to" header, CircleDot icons, current state marked by Check), and added HTML5 drag-and-drop on top — the target column highlights on drag-over, the source card goes semi-transparent during dragging. The mobile agent rail was implemented as overflow-x-auto no-scrollbar inside a strict overflow-x-hidden main container, so the author reported no whole-page horizontal overflow at 393px. These are choices in this output, not evidence that the model always behaves this way. A native select is not inherently a functional or accessibility defect; replacing it is a design tradeoff.

Second: mock data reads like a real project

Six agents carry personal names (Atlas, Pixel, Forge, Sentry, Lens, Mason) with distinct Claude model labels (opus-4.7 / sonnet-4.6 / haiku-4.5), illustrating a role-to-model UI concept; these are not actual routing or billing records. Task IDs run T-2041 to T-2049 plus three done-state tasks, file paths reach concrete strings like src/checkout/PaymentForm.tsx, risk numbers ("Coverage fell from 82.1% to 77.9% after PaymentForm rewrite — 3 new branches lack tests") line up with diff stats in the activity timeline ("Updated PaymentForm.tsx (+412 / −287)"). The internal consistency is a step beyond GPT-5.5's already-good run.

Third: the theming system is engineered, not hardcoded

Instead of a binary light / dark swap, every color is exposed as an RGB triplet (--bg-base: 247 248 250), so Tailwind can do bg-status-running/10 with proper alpha derivation. Status colors (running / idle / blocked / reviewing / done) live in the same variable layer — re-skinning the entire app is two :root blocks. The prompt didn't ask for this, but Claude went there anyway. Combined with a useTheme hook that reads prefers-color-scheme, persists to localStorage, and toggles a class="dark" on the html element, the author recorded no obvious switching flash, without a preserved first-render measurement.

Fourth: interaction polish above the demo bar

Drawer opens lock document.body.style.overflow, close on ESC and on backdrop click, ride animate-slideIn in; project dropdown listens to outside mousedown; risk dismiss appears only on hover so the main view stays calm; running-status dots get a pulseRing animation; KPI numbers turn on font-variant-numeric: tabular-nums for stable column alignment. Each item alone is small. Together they make the page feel like the engineer who built it actually ships product.

4. Where It Still Falls Short

Main chunk is slightly larger than the GPT-5.5 run

The historical main-JS figure is 599.24KB raw (170.40KB gzip), about 38KB above the earlier record of roughly 561KB, with a Vite Some chunks are larger than 500 kB warning. No bundle attribution report is available, so the difference cannot be assigned to agent count, chart modes, or dragging, much less model efficiency. Inspect dependency weight and loading paths before deciding whether charts or the drawer should be lazy-loaded.

Desktop dragging was observed; touch behavior still needs validation

The author recorded working desktop mouse dragging and task-state changes through the per-card “...” menu. The implementation used HTML5 native draggable, but no iOS Safari device or other touch-browser operation record is preserved here. Unverified behavior is not proof of browser incompatibility. Validate touch, keyboard, and menu alternatives before deciding whether a different drag library is needed.

No focus trap inside the drawer

Keyboard navigation works — buttons, roles, and aria attributes are present, and .focus-ring is defined as a utility class — but Tab can leak out of an open drawer back into the underlying page. Doesn't show up in casual demos, but it warrants a focus-management and complete keyboard-flow review. It's the kind of "easy-to-add but skipped" item that separates a demo from a shipped feature.

Full-page Pixel 5 viewport capture (393px): a top agent rail, two-column KPIs, and a vertical board. The author reported no whole-page horizontal overflow, but the screenshot alone cannot verify rail scrolling, touch dragging, or access to hidden content.

5. Same Follow-up Interaction Test, This Run

In the GPT-5.5 review I ran a combined operation: move a backlog task to In Progress, switch to 7 Days, then filter to high risk only. I ran the same combo here. The kanban accepted both drag and the "..." menu — during drag the source card went semi-transparent and the target column showed an accent border. The 7-day switch recomputed the Active Agents scale, Estimated Cost, and token-usage chart together. Severity pills allow multi-select and full deselect; if you turn all three off, the filter resets back to "all on" so the list never collapses to empty.

Reading the source, those state slices live at the App level via useState, with useMemo deriving series / throughput / metrics and filteredRisks. Like GPT-5.5, no Redux or Zustand. Unlike GPT-5.5, the line between source data and derived data is sharper here: kanban tasks are mutated through setTasks, risk dismissals through setRisks, severity filtering only touches the UI layer. That clarity is exactly what makes the difference between a demo that can plug into a real backend and one that has to be rewritten for it.

The boundary, again, is honest: refresh resets everything, the Pause / Reassign buttons in the drawer carry no handlers, and the activity timeline does not gain a new entry when I drag a task. "Internal state coupling" was implemented; "operation logging" and "persistence" were left for the next phase.

6. Same Prompt, Two Deliveries — Side by Side

The author’s two recorded outputs show the following differences. Checklist counts are not model-capability scores.

  • GPT-5.5: Eight recorded demo checks met, one partial (mobile reachability unverified). Status changes via native select. 5 agents, 1 project. Main chunk 561KB.
  • Claude Code: Nine recorded display checks met. Status changes via custom menu + drag. 6 agents with mock model labels. 3 switchable projects. Main chunk 599KB (+38KB).

Both recorded outputs were useful for a demo discussion. In this pair, I preferred the Claude Code output’s custom state menu, narrow layout, and theme organization; the Codex output had the smaller recorded main bundle. A native select is not an error, and custom controls are not automatically more accessible. These are prototype-specific tradeoffs.

I would use these prototypes to discuss information structure and flows, then separately review persistence, logs, focus management, and real interfaces. A preference between this pair should not become a recommendation for every front-end project.

As a personal workflow preference, I often start code search and debugging with Codex and try a first layout with Claude Code. That preference was not systematically sampled and is not a capability split established by this demo. Validate the choice in your own repository, configuration, and tasks.

The detailed Chinese brief fixed many product and interaction decisions and may have reduced the room for improvisation. There was no short-prompt comparison or repeated run, so neither a percentage contribution nor the likely gap under a shorter prompt can be established.

Method and Result Questions

Why is “Claude Code 4.7” no longer treated as a confirmed version?

Claude Code is the execution tool; its foundation model needs a separate record. The old wording and URL retain a historical 4.7 label, but this page has no CLI version output, session model ID, or request log. The opus-4.7, sonnet-4.6, and haiku-4.5 labels in demo cards are mock data, not evidence of the model that generated the project.

Do nine checks mean the result is production-ready?

No. These were the author’s recorded display and basic-operation checks, not an automated pass rate. Drawer actions such as Pause/Reassign lack business handlers. Focus containment, persistence, operation logs, and real API integration remain unfinished; mobile evidence is limited to a 393px screenshot and the author’s observations.

Do 599KB versus 561KB establish model code efficiency?

No. They are historical main-JavaScript chunk figures, about 38KB apart. Dependency locks, bundle attribution reports, and repeated builds under matched conditions are not available here. Features and dependencies may affect size, but the total does not establish the cause or relative model efficiency.

Was mobile drag-and-drop verified?

No iOS Safari device-touch test record is available. The author reported desktop dragging and task-menu operations. The screenshot shows a narrow layout; it cannot verify touch dragging or menu focus behavior. Test dragging, menu alternatives, and keyboard operation on target devices.

Can this choose between Sonnet, Opus, or a 1M context window?

No. There was no independent Sonnet/Opus comparison or verifiable context configuration or token-usage record. The model cards are examples. Check current official configuration and actual session records for available models and context limits.

Does the same prompt make this a controlled experiment?

Not by itself. The author recorded using the same Chinese brief, but this page has an excerpt rather than full sessions, parameters, tool versions, and repeated runs. Compare only these two outputs. No short-prompt comparison was performed, so no percentage contribution can be assigned to prompt detail.