[2605.06869] Agentick: A Unified Benchmark for General Sequential

[2605.06869] Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

arXiv - AI May 11, 2026 4 min read

About this article

Abstract page for arXiv paper 2605.06869: Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

Computer Science > Artificial Intelligence arXiv:2605.06869 (cs) [Submitted on 7 May 2026] Title:Agentick: A Unified Benchmark for General Sequential Decision-Making Agents Authors:Roger Creus Castanyer, Pablo Samuel Castro, Glen Berseth View a PDF of the paper titled Agentick: A Unified Benchmark for General Sequential Decision-Making Agents, by Roger Creus Castanyer and 2 other authors View PDF HTML (experimental) Abstract:AI agent research spans a wide spectrum: from RL agents that learn from scratch to foundation model agents that leverage pre-trained knowledge, yet no unified benchmark enables fair comparison across these approaches. We present Agentick, a benchmark for sequential decision-making agents designed to evaluate RL, LLM, VLM, hybrid, and human agents on common ground and to power research on the fundamental challenges of sequential decision-making. Agentick provides 37 procedurally generated tasks across six capability categories, four difficulty levels, and five observation modalities, all exposed through a single Gymnasium-compatible interface. The benchmark ships with a Coding API, oracle reference policies for all tasks, pre-built SFT datasets, a composable agent harness, and a live leaderboard. An evaluation spanning 27 configurations and over 90,000 episodes reveals that no single approach dominates: GPT-5 mini leads overall at 0.309 oracle-normalized score while PPO dominates planning and multi-agent tasks; the reasoning harness multiplies LLM perfo...

Originally published on May 11, 2026. Curated by AI News.

Llms

Researchers asked ChatGPT, Gemini and Claude which jobs are most exposed to AI. The chatbots wildly diagree

A study reveals that AI models disagree on which jobs are most vulnerable to automation, highlighting the unreliability of AI-generated e...

AI Tools & Products · 4 min · about 6 hours ago

Llms

I stopped treating ChatGPT like Google — and everything suddenly clicked

I stopped using ChatGPT like Google and started treating it like a thinking partner — here’s why that simple shift made the AI dramatical...

AI Tools & Products · 8 min · about 6 hours ago

Llms

Hackers abuse Google ads, Claude.ai chats to push Mac malware

AI Tools & Products · 6 min · about 6 hours ago

Llms

Does Claude dream of electric gavels? A federal case with Kansas connections sets an AI precedent.

AI Tools & Products · about 6 hours ago

[2605.06869] Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

About this article

Related Articles

Researchers asked ChatGPT, Gemini and Claude which jobs are most exposed to AI. The chatbots wildly diagree

I stopped treating ChatGPT like Google — and everything suddenly clicked

Hackers abuse Google ads, Claude.ai chats to push Mac malware

Does Claude dream of electric gavels? A federal case with Kansas connections sets an AI precedent.

No comments

Stay updated with AI News