Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

How am I supposed to take this tool seriously if it struggles with the concept of "dont"?

if what you're suggesting is true, then more important instructions like "Don't delete the production database" are a problem. I shouldn't need to consider how to phrase "Don't delete the production database" in a positive manner. Isn't the point of an AI agent that it understands my intent and I don't need to hold its hand? I'm not saying your suggestion is wrong, I just think that suggests a limit to what these tools should be used for if that's the case.

My guess though is that it probably ignores some of the positive instructions too, and the hardcore AI users mostly don't notice because they probably aren't reviewing the work.



Absolutely. It should simply not have permissions to delete your production database, can’t rely on prompting alone for that kind of safety.

These are still stochastic machines, guardrails must be inserted at the system level.

They are getting better every day about managing their own guardrails, so we will get there eventually.


I entirely agree, although I think this is counter to the narrative that we should just let the agent do everything that the labs want to sell.

There's a really major problem with the concept of human-in-the-loop though, which is just that humans are not built for that that kind of work. It's like all those tests against the TSA where they manage to sneak something through that should have been caught. But the problem is, a TSA agent sees probably like 1000 things they can ignore to the 1 thing they need to look at, and it's easy to just start rubber stamping things.


If you're relying on asking the LLM "pwease don't delete" then you're already in trouble. This kind of stuff doesn't work with people either and they generally exhibit actual signs of intelligence.


Sure; it's an example, I would never rely "Please don't delete" for real important data. However, we're being sold an idea that we should just let the agent do everything, so I think it's important to point out how utterly insane that idea actually is.


> How am I supposed to take this tool seriously if it struggles with the concept of "dont"?

Don't think of a pink elephant in a tutu.

What did you just think about? =)

This has been true since the very first GPT release, any words you say to the model will nudge it towards that word - it doesn't matter if you have a negative in front of the word.

Form the guidance using positive words like "always use the development database at..." instead of "don't access the production database" ... Now it "knows" that the prod database exists and "thinks" about it when processing.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: