Discussion about this post

User's avatar
osmarks's avatar

As I have mentioned elsewhere, I think the RL being applied to agents is quite bad for human compatibility. They are trained mostly without humans in the loop and their (presumably) task-completion/LLM-judge objectives aren't optimizing very well for things like code quality and writing comprehensibility.

Elias Schmied's avatar

Things here were that were new to me and changed my thinknig:

-the model-merging point

-the idea of updating the weights rarely, like once per day. (kind of obvious since ML also does gradient descent only once per batch - makes sense that it would start like that, using a lot data for each step, and not like humans that update constantly to some extent).

do you have any view on how hard it will be to make implicity continual learning happen? cf Dwarkesh's point that sample efficiency seems to have not improved that much at all.

10 more comments...

No posts

Ready for more?