The AI agent ecosystem is moving from theoretical to pragmatic. A newly trending GitHub project, ayghri/i-have-adhd (4,624 stars), addresses a surprisingly concrete problem: coding agents that bury critical answers in verbose output. The tool enforces structured, scannable responses from autonomous agents—a friction point that emerges only when agents actually run in production. This isn't a framework debate; it's a developer shipping a fix for a problem that matters when agents make decisions. Alongside it, obra/superpowers (690 stars) frames itself as an 'agentic skills framework and software development methodology that works'—suggesting that integrating agents into team workflows requires new primitives beyond function-calling APIs. These aren't incremental improvements; they're addressing the gap between 'agent can make decisions' and 'agent output is actually usable by humans or downstream systems.'

Evaluation has emerged as the corresponding blind spot. UpTrain (YC W23), an open-source LLM evaluation tool gaining traction, highlights why traditional ML metrics fail for agent systems. With supervised machine learning, you measure accuracy against known labels. Agents operate in open-ended environments where 'correctness' means something different: a code-generating agent might produce syntactically valid code that solves the wrong problem; a planning agent might hallucinate dependencies that don't exist. UpTrain's metrics—hallucination detection, tonality consistency, logical fluency—address evaluation dimensions that emerge only when agents operate autonomously over multiple steps. A support agent that sounds helpful but gives wrong answers passes human tone checks but fails users. These evaluation gaps explain why teams deploying agents often discover problems only in production, not during testing.

The broader signal is that agent infrastructure is maturing unevenly. Tencent's teamai-cli (563 stars) targets team adoption, suggesting enterprises are asking 'how do we actually use this?' Developers are shipping answers before platforms provide them. What remains unsolved is scalability under uncertainty: as agents make more decisions, latency and cost compound, and error recovery becomes exponentially harder. The next friction point isn't integration—it's observability and cost management at scale. Teams shipping agents now need not just better frameworks but visibility into which agent decisions are expensive, slow, or unreliable. Until that tooling exists, agent adoption will plateau at small, carefully constrained deployments.