TL;DR: A test of Claude, Gemini, and o3 on a tree-based combobox found that LLMs can scaffold compound component APIs quickly, according to WorkOS, but they still struggle with nested behaviour, keyboard support, screen-reader semantics, and state coordination in complex UI. That makes context, tests, and manual review the real guardrails, not prompt length alone.
Editorial analysis by NHI Mgmt Group, based on content published by WorkOS: “Vibecoding a complex combobox component”.
Key questions
Q: How should teams use LLMs safely for complex UI components?
A: Use LLMs for scaffolding, boilerplate, and pattern completion, but require tests and human review for behaviour, accessibility, and state coordination.
Q: What breaks when a generated component looks correct but the interaction model is wrong?
A: The usual failure points are nested state, focus management, keyboard navigation, and accessibility semantics.
Q: How do security and platform teams know when to stop prompting and rewrite code manually?
A: Stop when repeated prompts are producing regressions, when stateful behaviour keeps drifting, or when testing shows that the tool understands the shape of the component but not the logic behind it.
Practitioner guidance
- Start with behaviour tests Codify keyboard paths, expansion rules, and screen-reader expectations before using generated code so the model has an explicit contract to satisfy.
- Embed component context in files Use comments, local examples, and design-system patterns inside the codebase so intent survives after chat context is lost.
- Review accessibility as runtime behaviour Test nested widgets with screen readers and keyboard-only flows, because correct-looking markup can still fail users in practice.
Bottom line: LLM-assisted UI generation can accelerate scaffolding, but it does not reliably preserve the behaviour of complex nested components.
Explore further
View Full Forum → | NHI Foundation Course → | Our Services → | Read the full analysis →
LLM-assisted UI generation is a scaffolding accelerator, not a behaviour guarantee. WorkOS’s example shows that the first pass can be structurally close while still failing on the hard parts: nested state, keyboard interaction, and accessibility semantics. That distinction matters across identity and admin tooling, where the component may look complete but still fail under real user behaviour. The practitioner conclusion is that generated code should be judged on interaction fidelity, not visual completeness.
A few things that frame the scale:
- AI-related credential leaks surged 81.5% year-over-year in 2025, with the surrounding AI infrastructure leaking 5x faster than core LLM providers, according to the State of Secrets Sprawl 2026.
A question worth separating out:
Q: Why do complex, nested controls need more than code generation to be reliable?
A: Because nested controls combine search, hierarchy, selection, and accessibility into one state machine. Generating the surface API is relatively easy, but maintaining correct behaviour across all those layers requires design intent, tests, and careful review.
👉 Read our full editorial: LLMs can scaffold complex UI, but accessibility still breaks