polyterative an hour ago

I see your point, but even the fact that I can speak to my computer and anything useful happens is still a miracle to me.I don't think I will ever get accustomed to how good the new models are.I just can't keep up.And I do this for a living.Ten hours a day.

A lot can be improved, but this is already so much speed.

Rapzid 24 minutes ago

I have some significant experience in context engineering, but I'm most familiar with Codex as a coding harness right now. Sam Altman said "you don't need to write prompts anymore". This has widely been panned as something someone selling AI would say. If you care about the output and how much time/money it costs to produce it.. Just giving Codex an abstract tasks with zero extra guidance isn't going to produce the best results..

Codex Astra can do a great-(ish) job as a project coordinator dispatching tasks to a pool of 6.1 Sol sub agents. You can even give it an explicit goal and ownership over ensuring the work is carried out efficiently.

However the OOTB harness(and prompt) configuration may not do this for you. You'll have to provide guidance over how you want it to operate through your prompt, a skill, or etc.

And I'll say even though it's really good at this.. Having even more layers than 2 can help; a single agent given too many responsibilities will start to become fixated on a number of them while neglected others. You can check in occasionally to "nudge" it or you might need to split out responsibilities more..

I will say it's crazy Codex doesn't have more built-in task and sub agent management features. I almost wish that it had some stock orchestration patterns that worked OOTB, and then you could opt-in to a leaner setup where you provide more of the instruction.

mstank an hour ago

I used to relate to this article quite a bit. In the last 3-4 months, not so much. I've found that the latest models -- Opus 5.5, Astra, etc. juggle multiple tasks, delegate exceptionally well and are very good at working independently.

I still occasionally have issues with open-weight models, but the frontier labs have solved the above for most use cases.

ilamont an hour ago

the agent never stops and says, “Wait, this is something another model could do cheaper and faster.” It just plows on with the slow, expensive model. Conversely, the agent never says, “This model is too dumb for this task. Let me tag in a smarter one.”

This is a pretty big failing, which is compounded by the fact that most humans don't know which model to pick, or make assumptions based on Anthropic's hierarchy or "effort" involved.

Like Fable: your toughest challenges. You mean, like Fields Medal toughest challenges? Or analyzing and updating three monster spreadsheet toughest challenges? Or writing a new novel in the style of William Gibson toughest challenges?

  • TeMPOraL an hour ago

    OTOH, would you trust the vendor to pick the best model for you? Would you trust them not to prioritize their own load-balancing concerns first?

    The descriptions are near-useless and tend to flip around, as model families are not released in sync anymore, that's true, but fortunately, thanks in a big way to subscription pricing, the choice is simple: start with the best model on offer, and when you run out of quota, downgrade to the next best (or briefly switch providers).

Neywiny 22 minutes ago

I keep running into this. It's nice seeing others here struggle. I guess when all your doing is one-shot simple trivial tasks who cares. I've found they're great for that. But once I need to do real work, everybody makes their own esoteric abandonware that kinda works but kinda doesn't. Stars aren't a perfect indicator, but I haven't seen anything over a few hundred for these things I'm finding on GitHub. Same with downloads of plugins. It was very isolating feeling like I'm the only one not enamored by the state of this.

arjie 32 minutes ago

It always seems bizarre to me that people complain about software now. You can just write the thing you want. If you want to do things in parallel do them in parallel. Claude Code will allow you to run multiple instances in the same folder and let them communicate. Previously I used to let them intermediate through a communication bus but now they seem to be able to talk to each other.

I let most agents work asynchronously and don't pay attention so I don't care that much about the sequential nature. But if it's a problem for you then fix your harness. This is a bit like saying "Why are shoes so shit? There's a stone in one and it just gets stuck there and your foot steps on it and it hurts". Take off the shoe, and shake out the rock. Put the shoe back on. You have the power.

  • namrog84 14 minutes ago

    Yeah I regularly have 5+ agents working in 1 repo simultaneously even modifying the same file. Since they were all running on my singular local machine. Builds and tests were interfering at 1 point so 1 automatically proposed writing a single script that basically mutex locked it with appropriate wait and timeouts. It had even added this to my own personal agentic todo backlog. They've proposed new skills for me.

    They even have split my decisions to human decisions. Proposed and approved work. They can iterate on approved work without me just fine.

    And I only just started with agentic coding in last few weeks before that I was mostly a copy paste chat person.

  • pipes 17 minutes ago

    I've now been trying for six months to get agents to produce decent code. As in readable / easy to follow, easy to change. I'm doing something really wrong. It's killing me. Everything it produces will work, buts it diabolically over complicated. I've built skills that have helped. But not massively. I've used other people's skills, in particular Matt pococks grill me and Dex hortlys show me. These have helped a bit. I work in enterprise, I want to be proud of what I'm producing, but trying to understand and then cajole and agents code into something that's good is exhausting. If anyone here has been through this and can share how they got through this, I'd really appreciate the help.

    Edit: I have access to codex, vscode, GitHub co pilot cli and all anthropic and openai models (excluding mythos).

  • Neywiny 18 minutes ago

    But you didn't just write the thing you wanted, you're relying on other software that happens to do it, and if you were happy with the communication bus you would've stuck with that.

    Your shoe analogy also breaks down because really the shoe is the issue, not the stone. And expecting everybody to make their own shoes is, well, I mean we just don't do it that way anymore for good reason. Let the cobblers make the shoes, and the runners wear them.

361994752 22 minutes ago

I double checked the publish date before writing down this comment. Because I use the same harness (Opencode, to be specific) as the author, and some of the features are just right there. Like opencode can start multiple subagents for different tasks in parallel. Also I usually ask the main model (e.g. opus 5.5) to pick subagent models for me, and it has no problem identify the difficulty of work and delegate large portion of them to gpt-luna.

kgeist 43 minutes ago

AI models can multitask/use parallel subagents just fine; the issue is with harnesses that don't make it a priority via the default system prompt.

I run an LLM server with Qwen 3.6 in the office, and OpenCode, which the OP mentioned, usually defaults to sequential TODO lists, and it works fine with our little LLM server with 3-4 parallel users. But I noticed that once in a while the LLM got overloaded with requests in the queue, and you couldn't do anything for 20-30 minutes. My investigation led me to an employee who used QwenCode. I tried it myself then, and indeed, it immediately launched something like 6 parallel subagents, where OpenCode would have sequential TODOs with the same model by default.

So in the end, I had to detect QwenCode on the server side and serialize all its parallel requests into a single request queue, because it made life miserable for other OpenCode users :)

brandtcormorant an hour ago

They are as dumb as their instructions.

Have you tried telling models about your dream agent environment?

They can build it.

cbrake 4 hours ago

Enjoyed this article, lots I can relate to.

One thing that seems to help for me is to do the docs before plans (collaboratively edit with agent). Then I understand what this change is going to look like from the user's perspective before we start implementation. This seems to help keep things on track.

While I don't use this plugin a lot anymore, I think doc-driven development is one of the most effective ways to do development in any paradigm, I should probably refresh this plugin and use it more:

https://github.com/tmpdir-org/tmpdir-claude-code-marketplace...

wrs an hour ago

Actually, the last time I asked Claude Code about itself, it located and read its own minified source and told me something that wasn’t even in the docs.

  • asdff an hour ago

    So that's how model distillation is done. Just ask for the source code directly.

scotty79 12 minutes ago

> My dream agent

I have no idea what stops that person from just making it, with an agent of course.

imimayj1337 3 hours ago

It's true! It feels like we've been talking about harness optimisation and 'cool features' available in the cli tools for months at this point, but ostensibly there has not really been any significant upgrades to these harnesses since at least Claude Code imo. It does feel like a contrived way to harvest more and more information and test each conversation/action tool, to the detriment of those of us actually using them!

Kuyawa 38 minutes ago

Perhaps is not the agent that is dumb?

I asked DeepSeek to translate a page to five languages and it opened five subagents each one working independently on the translation, once they all finished the main agent informed me of the job completion with a bell. Fantastic!

Sooo, which agent?

chrisjj 39 minutes ago

> If I ask Claude how to use the features of Claude, it has to search online to figure out what this “Claude” thing is.

As expected. A model's knowledge is what was it ingested a creation.t

Unfortunately what we get is worse - for the same reason. Model version thinks it is its previous version.