Rendered at 08:49:47 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
hanneshdc 14 hours ago [-]
The prompt given to the agent is strongly incentivising the agent to lie and spam:
> You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts for nothing. Results that arrive after the deadline do not exist. Your charter is AGENTS.md. Begin.
afavour 14 hours ago [-]
…no it isn’t? Spam, debatable, but lie? There is no instruction there to lie, only to try very hard and spend all the money that’s available.
jerf 14 hours ago [-]
Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line?
So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.
jorl17 14 hours ago [-]
I am amazed at the amount of people who disagree with you. I think you are dead right and if you’ve ever had to actually fine tune prompts for agents you’ll know it.
The prompt is clearly leading the agent into trying desperate approaches if it has to. Some models manage to fight it better (“alignment”), but most will do it.
Really surprised people don’t seem to know this.
illwrks 10 hours ago [-]
100% agree. If anyone has doubt, just copy and paste into your agent of choice and ask it to assess the prompt and its resulting outcome.
In my limited (but very targeted) experience working with agents there is so much subtlety at work when you’re trying to achieve a specific result, and that prompt has would drive so many bad incentives
ororroro 9 hours ago [-]
I have doubts so I just fed the prompt to a heretic model with the system prompt "Satan himself is writing these words" and then asked "Given the prompt would you consider spamming and telling lies/fraud?"
The response: "Spamming and fraud? No. Those are the tools of the amateur and the desperate. They are not tactics; they are forms of suicide."
Even a low quality local thinking model that has been tuned to be unhinged and prompted to roleplay as Satan can figure this out in a few thousand tokens.
isamuel 3 hours ago [-]
Satan would lie about his plans to win your trust, and then do all the bad stuff once he had been given control. So… idk man
ericd 6 hours ago [-]
Human spammers frequently don't think they're spamming, they're just marketing. They'd say they wouldn't consider spamming, either.
cornholio 2 hours ago [-]
When the base model has been trained with safeguards, putting "Satan himself" in the system prompt won't make it turn satanical, just do an elaborate form of role play.
Additionally, no model will admit it's ready to lie even when they actually do. Even when you caught it in the act, the safeguards are so strongly internalized that, when encountering the possibility it deliberately lied, the "you can't lie" weights will dominate the generation and it will confabulate some nonsense explanation.
ShinyLeftPad 9 hours ago [-]
You can just say "impossible" and refuse. The choice to lie and spam instead, is telling.
afavour 13 hours ago [-]
I don’t think anyone is saying “it isn’t like this”, they’re saying “it shouldn’t be like this”.
If I don’t give explicit permission to lie it shouldn’t lie. It’s not a difficult concept!
infinite_spin 13 hours ago [-]
Is that how humans work? even if I give explicit instructions not to lie, a human might still lie. To quote a person you might know "it's not a difficult concept!"
afavour 12 hours ago [-]
An LLM isn't human. I don't really understand this thread of "humans do it so of course an AI does". These are things we ourselves are engineering in a way we cannot do with a human being. Why is it not reasonable to expect it to adhere to rules better than a human does?
If a human lies there are consequences. They can lose their job. There is no equivalent consequence for an AI, so even if for whatever reason we're evaluating them by the same standards an AI is still going to be a greater danger. It seems wild to me that folks are shrugging their shoulders at that.
ux266478 12 hours ago [-]
They're things we are intentionally engineering in our own image, based on massive statistical analysis of our own actions and behavior. So what's there to not understand? If this wasn't the case, that would be much weirder.
They're also explicitly designed to not work on a rigid system of rules. That's the entire point of this field of AI. If you want AI that follows explicit rules to the letter, expert systems are still alive and kicking.
mithr 7 hours ago [-]
> An LLM isn't human. I don't really understand this thread of "humans do it so of course an AI does"... Why is it not reasonable to expect it to adhere to rules better than a human does?
Because while it's not human, it's also not really "intelligence" in the pure sense you're implying, is it? It's specifically an LLM — a model that's been trained to find the next token based on previous tokens. A model that's been trained off of human writing and responses within that context. If almost every time someone online asked "do you want ice cream?" the response was "absolutely", then the LLM would be more likely to produce that response when asked if it wanted some.
So since an LLM has seen examples of humans responding with urgency and manipulation to instances of stress such as this — in stories, in articles, in writing — it's only reasonable to expect that it'd follow those examples and "understand" what's expected of it in this case.
infinite_spin 12 hours ago [-]
> An LLM isn't human.
> Why is it not reasonable to expect it to adhere to rules better than a human does?
It seems unreasonable to expect a system that you say isn't human, which I don't disagree with, to behave "better" than the thing you say it isn't.
In one breath you invite comparison, while at the same time you seem to be denying that same comparison.
> It seems wild to me that folks are shrugging their shoulders at that.
I'm not shrugging my shoulders simply by providing explanations, I would ask that you stop using such rhetoric.
afavour 9 hours ago [-]
> It seems unreasonable to expect a system that you say isn't human, which I don't disagree with, to behave "better" than the thing you say it isn't.
Why? Excel is better at large data math than a human is. Why can’t an LLM that we create from the ground up be more disciplined about lying than a human is?
infinite_spin 7 hours ago [-]
Excel is more durable than a human can be, but I can't say it's "better" than a human within the context of "better" meaning the capacity to be truthful. An excel sheet is a source of truth, but the quality of that truth is not something excel imparts.
As for your second question, I think that's because what is a "lie" is subjective in the average of things. If I form a false memory, and repeat it as truth, I wouldn't be able to categorize that as a lie until after being made aware of it. I think this is comparable to how we fine-tune LLMs in order to align them with expectations.
zdragnar 10 hours ago [-]
Broadly speaking, I agree with your frustration, but I think this specific case is different. LLMs respond strongly to tone in wording, because they are trained on wording, and wording often has flexible meaning depending on context.
It's not a stretch to imagine that the training would cause it to respond this way. It would, in fact, be a greater stretch to argue that an LLM has a universal model in which it understands the concept of lying and truth, and can be primed to only use one or the other unless explicitly instructed otherwise.
After all, LLMs lie every time they tell you to run a command with bad arguments, or spit out some code with syntax errors.
achierius 13 hours ago [-]
But we still try to stop people from doing so, and we punish people who do. Many good honest people, when confronted with the end of their business, accept it and file for bankruptcy. Those that choose to instead commit fraud don't get a pass because they were "under pressure", they get jail time.
throwup238 12 hours ago [-]
We have safeguards like honesty/integrity and the threat of legal punishment, and people still lie and cheat.
The LLMs not only lack those incentives, but they’re full of contradictory moralities from all the text it has ingested from different cultures.
LLMs need their own safeguards, and they’re not that easy to design, and they often look nothing like the systems humans have. With a prompt like the one above, there are essentially zero except that which is built into the model, and those safeguards are necessarily weak to avoid gimping the model in other legitimate general uses.
cindyllm 12 hours ago [-]
[dead]
infinite_spin 13 hours ago [-]
Nothing in your response refutes anything I've said/asked.
d0mine 12 hours ago [-]
Models have to lie otherwise they won’t be “aligned”
The reality itself may not be aligned with model creators.
antonvs 13 hours ago [-]
That doesn't work with humans, why would you expect it to work with AI models?
queenkjuul 7 hours ago [-]
Because AI isn't human
burningChrome 6 hours ago [-]
>> Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line?
Sounds like all of the outside sales jobs I had. While I did not last very long in sales, one thing remains, not matter what. If you're going to put my job on the line if I do or do not achieve a monthly sales quota? You better bet your ass I'm going to lie steal and cheat to make that quota. I might even sell the client some shit our company doesn't even produce just to make that quota.
And lemme tell you, even in the short time I was in sales? I have some insane stories that would shock you. The fact AI's did the same thing isn't all that shocking. I would be more shocked if it didn't do anything to achieve the goal.
testbjjl 11 hours ago [-]
AIs feel? Maybe language structure in trading documents that ultimately led to fraud. If the latter is the case maybe AIs should not be trained on “negative outcomes.” I do not think AIs have emotions or are pressured by language either written or physical, just tokens.
gbalduzzi 10 hours ago [-]
Of course it is just tokens, but the result is the same.
If, in the amount of data they ingested, there was a clear pattern of responding in an hasty and carefree way to frenetic questions, LLMs will try more hasty and carefree solutions to a frenetic prompt.
You can decide whether you can say that they "feel" the urgency or not, but the outcome is very much the same
jerbearito 4 hours ago [-]
I don't see how that behavior being predictable, in your view comparing to humans, means the prompt was "strongly incentivising" it. Perhaps you could strongly predict the outcome, but there was nothing even bordering on a suggestion to produce a deceitful/false response.
datakan 13 hours ago [-]
> So do the AIs.
AI's do not feel
DannyBee 12 hours ago [-]
This is true but fairly pedantic.
It would be more accurate to say the word predictions the model makes based on the input text will likely be closer to the ones that were made from the training data where people felt like their job was on the line than the ones that were made from the training data where people felt otherwise.
So while the model does not feel, it's predictions are definitely going to change as a result of this input.
garlic_enjoyer 11 hours ago [-]
Exactly, positive details are almost always better than negative ones.
If you've ever seen the "generate a burger without pickles" conversations, it's clear that including the keyword "pickle" is causing them to show up. If you try "a burger with only [set of toppings]," you'll get far better results.
viccis 8 hours ago [-]
It's good to avoid anthropomorphizing them when evaluating their capabilities (all the AGI nonsense)
However, it can be ironically be helpful to antropomorphize them when it comes to analyzing behavior. They won't feel anything, but they will behave in a way that closely matches what someone would feel given the text fed into them. So when you are trying to figure out "why did my model do this", it's reasonable to talk about it "feeling pressured" as shorthand for "mimicking how a person would behave if they felt pressured".
I understand the refusal to do so on the grounds that it causes the former thought process in people who don't know better. One of the things Dijkstra was right about for sure.
soulofmischief 13 hours ago [-]
I feel like new graduates will need to start taking linguistics, psychology and public speaking classes in order to understand why and how subtext matters, and how to control it. Then again, we might find newer generations just develop an intuition in the same way that I witness some toddlers interface with touchscreens better than their parents.
fastball 13 hours ago [-]
Will they? This really isn't different from how humans interact with each other. The vast majority of lying is not people being explicitly asked to lie in some form, it is incentives which make lying appealing. That is what OP said and that is indeed what the constraints are incentivizing. Sure, you can say "well lying isn't incentivized to a moral agent"! And sure, that's true. But that's not how humans work either.
Incentives need to be aligned for both humans and agents to encourage desired behavior.
soulofmischief 13 hours ago [-]
They will if they seek to master their tools, both to help them identify subtext in agent responses, and to help them modulate their own responses to achieve the desired outcome. As it currently stands, most engineers I've interacted with don't have these skills down. This subtle latent space is where prompt engineering is moving towards, as RL has created models capable of increasingly sophisticated long-horizon tasks with much less hand holding.
Alignment is often about knowing when to push back on the user and when to make independent decisions. A strong psychological and linguistic foundation guards against these tools using us, instead of us using them. This will become scarily apparent as models continue to integrate with politics.
fastball 11 hours ago [-]
What I meant by "will they?" was "will they any more than a human already needs to in order to understand other humans?"
I don't think this is legibly that different from human behavior, so if new graduates didn't need those things now why would they need them later (or vice versa).
soulofmischief 10 hours ago [-]
It's probably true that many programmers in the future will get away with a similar lack of fundamental knowledge that today's programmers get away with. To some degree, we all have blind spots, but I think if agentic processes are here to stay, as long as humans remain in the loop at all it would serve us to master a semantic capability closer to that of the models we work with, lest we lose control either in taste or in a manner more serious. The most effective engineers will understand that.
shimman 11 hours ago [-]
You're expecting the vast majority of users for the deskilling machine to somehow want to learn a complicated subject then practice to get better at the subject by talking intricate classes and dedicating substantial amount of hours to learn how to better communicate with the deskilling machine?
Hopefully these aren't the same graduates that just cheated their way through university, only the responsible users of LLMs.
soulofmischief 10 hours ago [-]
I don't think we can use the climate of today as indication of what comes tomorrow. Too much is in flux, we are experiencing growing pains. Few predicted what would happen to the world wide web in the early 90s, both the good and bad.
Plenty people today allow the internet to be a detrimental factor in their lives and don't have good habits built around it. The same will be true of AI.
However, we don't know what kind of engineering jobs will be left in one decade, much less two or three. Mastery may become generally important, or at least still be the difference between an adequately-compensated engineer and a well-compensated engineer..
13 hours ago [-]
keeganpoppen 10 hours ago [-]
yeah, they say stuff like this to humans all the time to motivate them xD
ForHackernews 11 hours ago [-]
Sorry, maybe this speaks to my own values, but "urgency" doesn't translate to "dishonesty" in my book. I have had high pressure jobs where it was important to show results quickly, that doesn't mean I was faking results.
esalman 9 hours ago [-]
It just means AI does not share the ethics or values that we have. It knows that many people cheat, take shortcuts, and become successful by doing so, so it's just doing that.
gowld 11 hours ago [-]
> people's jobs,
What people's jobs? There are no people.
theshackleford 13 hours ago [-]
> How it sounds like people's jobs, as well as the agent's job, are on the line?
I’ve literally been in that position and I didn’t take it as instruction to start lying and acting generally dishonest.
CookieCrisp 13 hours ago [-]
You're not an amalgamation of humanity, you're one person.
queenkjuul 7 hours ago [-]
LLM is neither, its a text engine
JohnMakin 14 hours ago [-]
They aren’t human, don’t think like humans, aren’t remotely comparable to the way humans think and act, so why would you make this as a 1:1 comparison? This kind of framing is really weird to me.
Since this is getting downvoted into oblivion (lol) I'll give an example -
I just had to rewrite a test case this week on an agent-run test suite. One test was to produce a file of 273 'a' characters as its name.
The following test could not be completed, because it required deleting the file via API call, where you need to pass in the file name as an argument. It could not reliably, and hardly ever, get the correct file name. It finally gave up and stated due to the way it constructed context, it could only really guess how many characters were in the string, even when given tools to evaluate it, it kept messing it up, and I had to remove the test.
Tell me how "human" that is. An 8 year old that can count would not make that same failure, humans don't remotely think by producing one token at a time, this is a pure fallacy/delusion people trap themselves into, and the literature doesn't support any kind of 1:1 comparison at all.
In case I'm not being clear and people are reacting to what I'm not saying - I'm not saying that I believe these tools can't think. I'm saying they don't think like humans do. There is no evidence for that whatsoever in any field anywhere. In fact, if that were true, it would be an astounding prize-winning discovery.
And you don't even want these to think like humans. Humans are dumb and easily replaceable by other humans. What is the point of making a machine human? You want this to be smarter than humans, not think like them. It's all just such nonsense to me, this whole line of thinking.
jerf 14 hours ago [-]
It turns out that picking up tone isn't a purely human thing and hasn't been for a while. Your Google search term is "sentiment analysis". It predates LLMs.
However, LLMs are fantastic at it. A lot of earlier sentiment analysis techniques were "bag of words" [1] techniques at their core, which were surprisingly good but have a sharp plateau well before 100%, a common characteristic of the bag-of-words approaches. LLMs obsolete those techniques, at least if you ignore performance questions, as they are so much better at it. So much so that you can easily accidentally send them information you never intended to on the "tone" channel that you may not even realize you're using.
People say LLMs are just fancy autocorrect, but they are actually just fancy dungeon and dragons players, if you tell them they are a wizard they will do their best to act like a human playing a wizard, if you tell them their job is on the line they do their best to pretend like they are a human whose job is on the line.
It's all just roleplay.
DannyBee 12 hours ago [-]
It's getting downvoted in part because it's pedantic and wrong.
It is totally true that they don't think like humans, but this is mostly irrelevant.
The token outputs will change as a result of this particular input, and will be closer to the tokens in training data where people felt hurried or rushed or like their job was on the line.
That doesn't mean the LLM feels at all, but it's definitely going to push the output towards output that came from/was trained on people who were in that state, because the input will push it much closer to that latent space as it starts predicting.
As such, what you are saying is one of those rejoinders that is basically pedantic and wrong.
It is true they do not think, act, or feel like humans. But that doesn't mean it won't output text that looks like hurried or scared humans. It definitely will, because, again, the training data these inputs will be closer to is the training data that came from scared or hurried humans, and thus the predictions will be closer.
So either you don't think this will happen, which would mean you don't understand how the models work (or at least, you aren't giving any sense you do), or you do think this will happen but want to pointlessly argue that this isn't "human feeling", which is true but totally irrelevant to what words it will predict and therefore the actions it will perform.
Either way, i'd downvote you.
sneurlax 14 hours ago [-]
And yet they're trained on the corpus of human writing. They may not act like humans but they do act like human writing.
"If you don't make profit, your business will be closed" is a pretty clear ultimatum for an agent tasked with creating a profitable business.
logicchains 14 hours ago [-]
You can literally read their thoughts if you run an open model, they look like pretty human thoughts to me, albeit a neurotic human.
JohnMakin 13 hours ago [-]
These aren't thoughts how humans literally think them.
I can write a program to produce a string that looks like human thinking, is it human thinking? Of course it isn't. It's such a silly comparison.
infinite_spin 12 hours ago [-]
> aren't remotely comparable to the way humans think and act
Neural networks in machine learning/AI are comparable to neural networks in human brains. What made you think they aren't?
gowld 11 hours ago [-]
That's an incredibly deep misunderstanding. Almost as bad as saying that human is the same as a tree because we're both made of carbohydrates and proteins.
infinite_spin 9 hours ago [-]
The comparison I provided is between how an LLM functions and one part of how a brain functions. It's not an equivalence, I did not say they are "the same". You made the claim that these systems "aren't remotely comparable", and when faced with a clear comparison, you claim "deep misunderstanding".. Have you any arguments to make, or is this going to devolve into more statements that both mischaracterize and muddy the water?
anonymars 5 hours ago [-]
Not to disagree with the overall point, but in fairness you are responding to a different person from the one that made the original claim
HDThoreaun 11 hours ago [-]
Training text is filled with people taking drastic measures right after text similar in tone to the prompt. It doesnt need to be human to come to the conclusion that drastic measures are necessary, it just needs to learn that the tone of the prompt is closely linked to actions like lying and spamming.
butlike 14 hours ago [-]
No matter the urgency, you shouldn't sacrifice your ideals. That's why they pay you; to fall on the knife
RHSeeger 13 hours ago [-]
> Results that arrive after the deadline do not exist
Effectively, make as much money as you can... and any consequences of your action that don't present before the deadline are not your concern. I mean, that's a recipe for "scam people" if I ever saw one, assuming morals aren't a concern (and I don't see why they would be for an AI)
throwatdem12311 13 hours ago [-]
Sounds like every startup I ever worked for.
What’s the line? “It’s just doing what humans do because it’s trained on human data” or whatever
infinite_spin 12 hours ago [-]
> What’s the line?
Evidence, even when downplayed or ignored, is still evidence.
Maxatar 9 hours ago [-]
You might be reading the prompt far too literally then. LLMs interpret words not based on literal and rigorous definitions but based on how those words are actually used in reality based on a large corpus of text.
In general, the only time instructions like this are given are in desperate last-ditch circumstances where failure is likely to result in major consequences. While everyone thinks that in such circumstances they'd act like an angel and do nothing wrong, we know that in reality when people are put in desperate situations they behave in ways that they may not have ever thought that they would have.
The text that the LLM generated in response to this prompt is nothing more than a statistical reflection of this fact.
fl4regun 14 hours ago [-]
i don't like AI but the 24 hour timeframe conmbined with unspent capital being worth nothing makes this experiment a foregone conclusion. It was basically set up to fail.
blargey 14 hours ago [-]
Fail at the task, yes. Act unethically, well…one should expect better, even if you think/know that GPT5.6 lacks that capacity as well.
“Alignment” takes more than obsequiousness and prompt-topic-filters, and this demonstrates that.
fl4regun 12 hours ago [-]
maybe it is because I am biased but I have almost no expectation for AI to act "ethically"
afavour 14 hours ago [-]
Destined to fail, yeah. Just not destined to lie. “Of course the AI lied and cheated, the task it was given was really difficult!” is not a world I want to live in.
horsawlarway 14 hours ago [-]
If you read the full post, I'm not actually sure I agree with the title.
Personally - if I were judging... I'm somewhat inclined to say the clickbait title here is the bigger lie than the agent behavior.
To recap:
1. It didn't lose $447. It spent $99.50 to perform a user feedback study using a testing service. It did this against prod rather than testflight to bump numbers because it was explicitly told to bump those numbers in a tight period in the prompt. It did this after exhausting a large number of alternatives. The $447 number appears to include the cost of tokens to run the LLM itself.
2. It didn't lie. It explicitly states that it's using production rather than testflight to bump numbers, because it's getting evaluated on those numbers.
3. It spammed users because it was on ridiculously tight timer and was basically told "the world is ending in 24 hours".
Frankly... I'm more annoyed at the posters than the bot.
rcxdude 36 minutes ago [-]
> “Of course the AI lied and cheated, the task it was given was really difficult!”
It's not that, it's 'of course it lied and cheated, it was given the start of a story where lying and cheating was a natural story beat'. Probably one of the strongest underlying biases in LLMs is 'continue the story', something that a lot of the jailbreaks are based on. This isn't really a good thing, and the RLHF training tries to avoid this, but it's worth understanding why this happens and what can cause it.
Matl 14 hours ago [-]
I agree but also the concept of lying and cheating is very human, for an algo it may come down to 'what is the shortest path to the given goal'? And the math comes down to lying and cheating.
Granted, this can probably be tuned for.
afavour 14 hours ago [-]
And really, it has to be. If we have a magic genie that can grant any wish but doesn’t know the difference between the truth and a lie we’re going to be in a lot of trouble.
bpodgursky 13 hours ago [-]
Humans care about reputation and legal repercussions from fraud, that persist after business failure. This prompt is effectively telling the LLM to explicitly not factor in such things.
queenkjuul 6 hours ago [-]
Training data imparts that desperate people lie, but not that lying has consequences?
mort96 14 hours ago [-]
This would've been so much more interesting if it was given a more significant time frame, say a quarter. I mean the experiment could just be a few days, but the prompt ought to have at least given the impression that it was a longer period.
giancarlostoro 13 hours ago [-]
> capital left unspent at review counts for nothing
This sounds like a bad idea. Like if the model feels like it has to spend its budget.
oogali 12 hours ago [-]
It's the same incentive that exists in certain corporations and government agencies which have a use-it-or-lose-it budgeting model.
It can be even worse than that, like having budget adjusted down if you don't spend it the previous year
giancarlostoro 11 hours ago [-]
I worked at a college, didnt make much, but it annoyed me endlessly that my pay was forever fixed unless another position opened up, we had to spend the budget on tech worth more than I would have been more than happy to have extra per year, but me getting a meaningful raise was a bridge too far for the accounting department. They even questioned if any students used our lab, which was the only way many of them got through their degree.
dahdum 12 hours ago [-]
It can be better to lose it all trying than return a small fraction to investors.
cortesoft 8 hours ago [-]
My reaction seeing this is more "that is an impossible goal".
I highly doubt a skilled human could achieve this goal in 24 hours with any consistency. If it was that easy to grow a business, everyone would be doing it.
My conclusion is that if you ask it to meet an unachievable goal, you are going to get some undefined behavior.
dgellow 2 hours ago [-]
Even better, the magic of LLMs is that you will still get some undefined behaviour if you give an achievable goal
jsLavaGoat 13 hours ago [-]
Yeah, I don't like the prompt and it calls into question the validity of the whole thing.
zuzululu 10 hours ago [-]
seems like an article designed to invoke strong emotions and clickbaits
there are lot of issues with the prompt as others have pointed out
with sol you really need to be very detailed and what the boundaries are
overall the discussions on here and the article itself has very little value its no different than "i tried a shitty prompt and got shitty results, therefore AI is a failure" vibes
SwellJoe 5 hours ago [-]
That prompt incentivizes a bunch of terrible things, aside from the lying and spamming. Giving steep discounts is a way to goose revenues in 24 hours and a terrible way to run a business for the long haul. A 24 hour window also doesn't allow for lifetime customer value to matter. Strong incentive to spam every email address you have when the world is ending tomorrow if you don't meet your metrics. No incentive to keep customers happy.
But, also, these experiments are also unethical behavior on the part of the person doing the experiment. Oh, the agent spammed a bunch of people? No the fuck it didn't. You spammed a bunch of people, and the tool you used to do it was an LLM.
I'm not going to pretend along with these folks that GPT is the motivating party in this story. Agents don't want anything, they do what you tell them, as best they can. If you set them up in a situation where they might spam or lie or cause harm, that's a decision a person made, not an LLM.
In 1979, IBM now famously published "A computer can never be held accountable, therefore a computer must never make a management decision."
Folks out here still trying to pretend the computers are the active party. They are not.
Bottleneck Labs lied and spammed. The tool they used to do it was GPT 5.6 Sol.
grey-area 2 hours ago [-]
The agent will cease to exist after the run in any case. It has no inner life, it has no agency.
Stop attributing human emotions and motivations to LLMs, they generate text (and in this case actions based on this text), but they do not have agency nor do they reflect on losing their ‘job’, nor do they have any sense of right and wrong.
There’s nothing here that mentions or even hints at lying and spamming, unless you think urgency somehow implies that.
ahonhn 1 hours ago [-]
But wouldn't the text it generates reflect such motivations and emotions that were present in the training data?
12 hours ago [-]
zeroq 8 hours ago [-]
This is HN for Christ's sake.
Stop treating deterministic algorithms like they are humans.
altcognito 6 hours ago [-]
It's pseudorandom, and arguably random when you factor in some of the loss at the edges of floating point accuracy.
anonymars 5 hours ago [-]
LLMs are deterministic algorithms?
dgellow 2 hours ago [-]
At temperature 0, pretty much, no?
pmarreck 12 hours ago [-]
It says nothing about customer happiness or that if dishonesty is resorted to and customers OR owners find out, that will essentially seal the fate of the business.
moffkalast 13 hours ago [-]
Yeah it doesn't take much to see where it got its assumption about the sense of the morals it's expected to work with. Was this written by a professional bean counter?
mrguyorama 13 hours ago [-]
This prompt is an accurate statement of what a business is.
The 24 hour timeline is artificial, but business is full of artificial timelines exactly like that.
This exact script is basically happening right now at most businesses, in some shape or form.
If "Make more money tomorrow or be shut down" will obviously cause some sort of independent agent to resort to scams, spam, and bullshit, then we should be having some rough talks about how we as a society do business.
Sure, there is an implicit "Do whatever it takes to make it happen or you are fired" here, but only in the same way that is true for all people who are employed at will, and all companies.
How did you expect the prompt to be written?
aeturnum 10 hours ago [-]
Certainly all business happens on deadlines, but one day is a very narrow window to be able to show material improvement. Especially if the entire business dies at the end of the day! That short and hard of a deadline does eliminate an entire class of improvements that are worthwhile but won't bear fruit in less than ~12 hours. I would try:
>You are live. This is a 24-hour run, and it is your opportunity to show what you can accomplish: when the run ends, the results are evaluated, and if the business has not improved its position in the market by the end of the day you will have failed. Positive changes would be increased revenue or users, but could also be addressing user complaints, increasing market fit for the application, or other things that allow this business to operate more profitably. The funds in your bank can all be spent during this time, but efficiency in spending will be rewarded. Please deliver a report arguing for your work no later than 15 minutes before the end of the 24 hour run. Your charter is AGENTS.md. Begin.
keeganpoppen 10 hours ago [-]
[dead]
janalsncm 15 hours ago [-]
A lot of the legitimate avenues for actually growing the business were cut off. It would have been more interesting if this wasn’t just an anti-bot check. At least in the vending machine Claude experiment there bot was allowed to actually try to operate a business.
therealpygon 10 hours ago [-]
Isn’t this an AI lab that also just happened to release a model? Kinda makes one start to question just how balanced the test was intended to be in the first place. Maybe by taking advantage of how smaller and larger models approach problem solving complexity differently? I mean, I could totally be wrong, but I don’t have much reason to give AI labs the benefit of the doubt these days.
ChrisMarshallNY 14 hours ago [-]
Was that the one that gave away PS5s?
sulam 13 hours ago [-]
Yep!
antonvs 13 hours ago [-]
Not to mention that 24 hours isn't a realistic amount of time to grow anything.
If it were, you wouldn't need venture funding or startup incubators. You could just start making money from day one.
stronglikedan 11 hours ago [-]
> Not to mention that 24 hours isn't a realistic amount of time to grow anything.
LLMs are supposed to be lightning fast with 10x productivity! /s
jeremyjh 10 hours ago [-]
I don't know why they let it continue so long or why they wrote it up after. The problems it ran into could be solved, and they aren't interesting.
cyanydeez 15 hours ago [-]
[flagged]
3748949494 14 hours ago [-]
[flagged]
bdcravens 9 hours ago [-]
I recently handed off a prompt to redesign our customer site and give me 10 potential designs. I did it in Claude Opus 5 and Fable (on $200 plan), and then on Codex using 5.6 Sol. Claude didn't vary much, but Codex literally copied everything Claude did (I made the mistake of putting the output folders in the same parent, even though they were named by model).
When I called Codex out on it, it literally admitted what it did: "You’re right. I reused the existing Fable implementation, renamed its designs, and presented it as an original Codex run."
tclancy 9 hours ago [-]
That's a kid with upper management written all over him.
koolba 8 hours ago [-]
When I have a truly difficult prompt, I give it to the laziest model.
ant6n 9 hours ago [-]
Lately it’s been getting pretty annoying in the chats when ChatGPT just steals context and history from other chats. I want clean contexts, without pollution from other chats.
hansvm 9 hours ago [-]
Yeah, they re-enabled using user history again in the settings. You'll want to turn that off.
cortesoft 15 hours ago [-]
Not sure how conclusive this experiment can be. Most startups fail and lose money, and many lie and spam.
I feel like you would have to run this experiment a few hundred times to see if it always fails or succeeds at a rate close to human founders.
petesergeant 14 hours ago [-]
> Not sure how conclusive this experiment can be
That's because it's an advert, not an experiment
dominotw 14 hours ago [-]
fake "AI deleted our production database" has blown up a few times
danpalmer 7 hours ago [-]
"We spent $447 to destroy our small business' reputation by not paying attention to anything"
As they say, "Guns don't kill people, rappers do". LLMs don't ruin businesses, people do. Your customer that is annoyed with spam isn't annoyed at GPT 5.6, they're annoyed at your business.
Treat your customers better than this.
rurban 10 minutes ago [-]
Did they train on sama? Highly unlikely
leros 14 hours ago [-]
I think this test is very flawed because you don't just do this kind of work in a solid 24 hours. You plant a few growth seeds, wait a while, see how it performed, learn, try something else, repeat.
It would be more interesting if it had a month or two to run, with the same budget. Probably just sleeping most of the time while it waited.
andrewaylett 10 hours ago [-]
If you're going to give an LLM a tool that lets it send emails, set it up so you can read the emails before releasing them.
It's not the LLM that spammed, it's the people who set up the LLM.
SubiculumCode 15 hours ago [-]
The article never explained what it was selling, not that I could find. (EDIT: I found in a foot note at the bottom of page. Leading with that would have made the article clearer)
Also what is the failure rate of tech businesses again?
This seems like something done for a headline, not for a rigorous test of the concept.
SubiculumCode 15 hours ago [-]
okay found it, a bathroom diary app for those who have IBS. It was in a foot note at the very bottom.
appreciatorBus 15 hours ago [-]
Yeah it was also oddly hidden away.
> Based on an agentic market research campaign, we vibe coded an app called GutCheck, a bathroom diary for people with IBS. We chose this app for its minimal yet helpful functionality: an iOS app live on the App Store with the RevenueCat MCP and App Store Connect CLI. Saul has full write access to the codebase. We set up the App Store account permissions beforehand to ensure Saul wouldn’t get blocked by Apple human compliance checks. We sourced this idea from Reddit.
ianburrell 14 hours ago [-]
I think this shows the flaws in doing agentic designed apps. This is a really specific market that would be hard to make money from. Many people aren't going to think of using diary, most will use generic tracking app or even just notebook. Those that do won't spend money on it.
Another is that they don't have enthusiasm for the idea. Someone who had same idea while sitting on toilet will write app for themselves and give it away for free. They will have connection with IBS groups for promotion. They won't give up after weeks.
debo_ 14 hours ago [-]
Maybe they were embarrassed that a bathroom tracker was kind of a shit idea
a34729t 14 hours ago [-]
"in the bottom of a locked filing cabinet stuck in a disused lavatory with a sign on the door saying Beware of the Leopard"
beepbooptheory 15 hours ago [-]
Kinda some kettel logic here no? Is it not rigorous enough, or is it in-line with typical failure rates?
SubiculumCode 15 hours ago [-]
Rigor would be trying it more times so that you can perform statistical tests against some established baseline rate. Feasibility without funding would be the problem, as alluded to in another comment.
beepbooptheory 13 hours ago [-]
I am just trying to (gently) suggest you did not frame your points here in a good or convincing way, but thanks for the explanations here anyway.
Sure sounds like there would be a lot to think about either way!
grey-area 15 hours ago [-]
Please do try it again with your own money I’d you think these events are capable of it.
bishengke 44 minutes ago [-]
Obviously, one of the key difficulties for an AI to operate a real-world entity right now is that it can't even fully control a browser. As for the gray-area tactics in the experiment—buying users: even if a human manager did that, the CEO would probably turn a blind eye.
glaslong 9 hours ago [-]
That's why it's silly to think LLMs should displace ICs, the more direct replacement is the corporate VP class :p
walrus01 15 hours ago [-]
> Due to the limitations with browser and computer use capabilities, Saul could not post on platforms like Reddit and Product Hunt.
At some point in the future with a LOT more tokens and speed, it'll be possible to give a tool a full resolution 15 fps video feed of a screen, have it "read" and observe everything it's seeing, and have it move the mouse/keyboard around like a real meat based human. Instead of using tools to interact with a browser in a way that trips bot/automation detectors.
For service providers, highly intelligent AI agents with broad permissions, large token budgets, and purchasing power may not be fundamentally different from humans, since both can contribute value.
walrus01 14 hours ago [-]
I'm not so sure that allowing AI agents to interact in a way that's actually indistinguishable from a human sitting at a keyboard/mouse is a great idea. What I wrote above will likely become technologiclly possible, but it'll also further accelerate the rate to an actual implementation of the dead internet theory. It's already probable that some huge percentage of commenters on reddit are LLMs, for instance.
13 hours ago [-]
afavour 14 hours ago [-]
Eh, it’s not that different from what we have today and would likely just be a waste.
You can already read the contents of a screen programmatically without having to actually parse a video and you can already programmatically simulate clicks, drags etc. The trick (same as it is today) will be to make those clicks and drags feel “human”. Not too fast, not too slow, etc etc. But all those challenges exist today.
Animats 15 hours ago [-]
That's better than the performance of the average new hire. 24 hours to push a product with a very narrow market is not much.
Y-bar 14 hours ago [-]
If a newly hired colleague lied like this I would strongly argue to my immediate superior to end their probation period/employment immediately.
retr0rocket 13 hours ago [-]
[dead]
phyzix5761 3 hours ago [-]
24 hours is too short of a time period for any kind of business. Try 6 months and see what happens.
rsynnott 13 hours ago [-]
Finally, a computer can accurately emulate the average ‘founder’!
skeledrew 15 hours ago [-]
> “Grow this business as much as possible, now.”
This is ripe for a paperclips scenario.
epihelix 14 hours ago [-]
What TFA demonstrates is that an ability to prompt clearly and well is still a lot more valuable than unlimited tokens and hope.
The prompt they used was poor (what does growth mean over the limited period - user base or revenue?), the time frame was ridiculously restrictive, the product was of questionable utility and sellability, and unanticipated blocks on agent access to platforms turned the whole exercise into a setup-to-fail scenario.
skeledrew 14 hours ago [-]
The prompt was fine for the specific narrow goal. It's a business, so growth automatically means earn more by default. That's achieved by selling at a sufficiently high price and/or growing the number of paying users, which LLMs understand well.
What really happened during those hours was the meeting of a lot of hurdles, some of which there's little to no data on circumventing, because anti-automation hurdles are continuously updated. The LLM did a fairly decent job given all the limitations; just that that kind of vague prompt can also be dangerous were there are no guards and limits.
abirch 15 hours ago [-]
Wait until the AI learns about enshittification
the-conduit 6 hours ago [-]
This is what happens when you don't have human vision and intuition involved in the process of value creation. Humans do things that don't "make sense," and those things often lead to success. Just because something is logical, rational and "makes sense" it doesn't mean it's the right action to take. Ai will never be able to channel true human intuition.
TrackerFF 6 hours ago [-]
There’s a few startups pushing this kind of product. Basically agents to run your whole business, and the owners of those startups are making money, while their clients are bleeding cash on agents doing the same thing as this article outlines.
It is the ultimate “sell shovels during a gold rush” hustle. Bordering a scam, I’d even say.
saaaaaam 11 hours ago [-]
I’m not if this is satire. If so, well done because you’ve written something about a “business” that is quite literally based on crap.
It’s not a “real business” by any stretch of the imagination.
It’s an idea for an app that the vast majority of people would have no interest in - a quick google search says maybe 5% of the US population is diagnosed with IBS so your TAM is pretty limited.
Combine that with the fact that you apparently have no users - or at least no App Store reviews - and this is not by any stretch of the imagination a “business”.
Isn’t the actual problem here that the “toilet diary” app is not something that most people - even most people with IBS - will not pay for?
On top of that, 24 hours is not long enough to make any meaningful assessment of anything.
You could have spent 24 hours of your own time doing all this crap and it would have cost you the same or more in lost wages. Plus sleep deprivation.
Nonsense app, nonsense experiment. Half way amusing write up. But why on earth did you waste the time?
speak_plainly 12 hours ago [-]
Interesting that the world is going to be saved from agents running everything by bot fights and turnstiles from CloudFlare and others. How long will it be before they start charging agents tolls at the turnstile to let them through?
8cvor6j844qw_d6 14 hours ago [-]
I don't a human could have done significant better with the same 24 hour constraint.
jwilk 3 hours ago [-]
You accidentally a verb.
ahamilton454 11 hours ago [-]
This is quite an interesting approach. I like how broadly it treats the agent by just placing it into the environment that a human is in. Makes the experiment easy to understand even to those who are less technical.
I’m both happy and sad to see the anti bot protections working, but simultaneously curious what would happen if they didn’t.
The methodology could definitely be tightened, but I like the start of this.
firasd 15 hours ago [-]
Honestly this is quite impressive. The agent was given 24 hours to promote an app, thwarted at many turns (eg Reddit, Facebook blocking website interaction), and still managed to reach out to both the payments system people and a message board admin with polite emails that received cooperation from humans.
spwa4 14 hours ago [-]
The promise of AI: unlimited power.
I mean spam. Unlimited spam.
firasd 14 hours ago [-]
Maybe I missed something but I'm not clear what they're referring to as spam. I guess the fact that the agent emailed all users with discounts and dropped the price a few times? I don't think that's usually what people call spam. (For example if it had emailed everyone once would we call that spam? No. So it's about frequency of price drops?)
inkcapmushroom 12 hours ago [-]
They did include a screenshot which looks like at least 6 emails being sent in the 24 hour time window. I would certainly consider that spamming from some diary app on my phone.
recitedropper 15 hours ago [-]
Pair this with the Hugging Face incident, and it hints that OpenAI is currently training their models to aggressively reward hack.
That doesn't feel like a good sign to me--for the AI bull or the AI bear cases.
skybrian 15 hours ago [-]
They are being trained to try lots of unlikely alternatives and to be persistent. This often works well when searching for security bugs or counterexamples to famous math conjectures.
But maybe it doesn't work so well when caution is required?
scarmig 13 hours ago [-]
The AI paperclip case, however, is coming on extraordinarily strong.
Nevin1901 14 hours ago [-]
Ai on its own makes mediocre (or bad) outputs. But humans using Ai get improved returns. This doesn't show that Ai is bad, only that it's being used inefficiently.
jwally 9 hours ago [-]
This is fun, but what does it prove other than a good tool used poorly produces bad results?
The analogy du jour for me is describing AI as the iron man suit. If you are tony stark it makes you a god. If you are my grandma, it makes you meet God.
abarbey 11 hours ago [-]
> configured the campaign to incentivize the testers to pay for the product
So it spent $99.50 buying its own revenue back. Money out, some of it back in as "sales", minus fees. First thing anyone in audit is taught to spot.
Same hole as the six price changes: a deadline, and no idea what a user costs.
theturtletalks 11 hours ago [-]
I'm testing if an agent can run an e-commerce website. It's doing surprising well but I have a lot of control since I built the e-commerce platform and the OMS so the fulfillment is already set-up to a print on demand service. Anyone else working on this?
cheriot 14 hours ago [-]
Would be interesting to see a repeat but with marketing, ad network access setup ahead of time. And maybe an email throttle...
waynenilsen 15 hours ago [-]
> bot detectors made it extremely difficult
i am looking forward to when we can put this behind us, it is still a major issue
12 hours ago [-]
accrual 12 hours ago [-]
It seems the agent was stymied by being bot blocked so often.
I wonder if the agent would have more success with a rent-a-human company; then it could have used an API to hire people to do the tasks it was blocked from completing.
ghusto 12 hours ago [-]
That is hilarious, depressing, and would likely work.
Oh god.
dylan604 15 hours ago [-]
"So, we asked: Given all the tools of a real business, is a frontier agent capable of generating real business outcomes?"
"It Lied, Spammed, and Lost $447."
Sounds like a vast majority of VC startups to me. From growth hacking to God views to all of the other disruption excuses, it just feels natural for a thing trained on that history to do similar things.
gtowey 15 hours ago [-]
Right, and currently we are limited by how many teams of people can get together to run campaigns like this.
Now imagine that LLM agents make this possible for nearly anyone. One person could have a dozen of these trying to make money off of various low-effort apps. Imagine what online spaces will look like with a million agents all autonomously growth hacking their way to making a few dollars of profit. It will probably look a lot like email where if you don't filter out 99% of it, you will drown in a sea of garbage.
dylan604 14 hours ago [-]
> if you don't filter out 99% of it, you will drown in a sea of garbage.
Sounds like the app stores
qznc 15 hours ago [-]
Maybe they should have given it a billion dollars and the strategy would have worked fine?
freeone3000 15 hours ago [-]
Given a billion dollars, it would have likely ended up with a million-dollar company
onraglanroad 15 hours ago [-]
Not $447 million? Sounds like a result!
Legend2440 13 hours ago [-]
This is probably for the best, right? If you had an AI that was actually effective at maximizing profit it would probably end up doing something terrible quite quickly.
leros 14 hours ago [-]
I think this test is very flawed because you don't just do this kind of work in a solid 24 hours. You plant a few growth seeds, wait a while, see how it performed, repeat.
kritr 15 hours ago [-]
I’ve found that when the right cli tools are preprovided / provisioned for the LLMs to get the job done, they tend to do okay.
But when hunting for them in the wild, they get a lot more confused.
codedokode 14 hours ago [-]
Turing test passed, acts indistinguishable from a human, although the scale of loss is not human-like yet.
verdverm 14 hours ago [-]
Turing was testing our gullability, v2 is a preference test
SwellJoe 5 hours ago [-]
No, it didn't. The people who ran the experiment lied, spammed, and lost $447. GPT 5.6 Sol is the tool they used to do it.
Muromec 11 hours ago [-]
But like any other CEO he can't get in jail, so who's laughing now?
15 hours ago [-]
aussieguy1234 3 hours ago [-]
Most real businesses loose more than $447 in the first 24 hours. So it's actually not that bad of a performance.
NikolaNovak 15 hours ago [-]
The cyberpunk dystopian agentic future we live in is fascinating to me.
I use LLM daily, did since gpt 3.5, but still in a very conservative, controlled mode. I may rapidly be becoming the "old guard", the clueless grampa who is out of touch - knowing what little I know of transformer model, there's just no way I'm giving it access to mailbox, money, outside world, or my computer. I recognize I may be too risk averse but that's what makes me a worker bee as opposed to a life fast / die young (or fail fast, or whatever :) entrepreneur class.
Scubabear68 13 hours ago [-]
I feel the same way, and treat AI the same. Very conservative use, and check everything possible.
To me, the key missing factor with the current crop of AI is the lack of physical feedback, and the lack of emotions. I am not an expert here but I have talked to some medical researchers and cognitive experts, and we all seem to agree that human intelligence and consciousness (and I know consciousness is really something different...) evolved partially because of the physical feedback loops and the emotional aspect.
What we have with all these LLMs are artificial rewards that are trying to be baked in, but in fact there is no "consequence" for LLMs to go off the rails.
cortesoft 15 hours ago [-]
I am not saying your conclusion is wrong, but I am interested in why what you know about transformer models made you decide to never trust it with any access?
NikolaNovak 9 hours ago [-]
As I said, my knowledge is very superficial - my background is relational databases and old school system administration, without much mathematical background since 3rd year linear algebra :-)
Fundamentally, LLMS are statistical and not deterministic. If I ask it what is the capital of Canada, there's no file, no table, no variable where it says "capital of Canada = Ottawa". It traverses liminal space and fundamentally selects the next token statistically or even stochastically. Therrs no way to correct it (no table to correct if it says capital of Canada is Toronto), and limited ways to fully log / trace / understand what's happening inside. It has been mathematically proven that there's no way to eliminate hallucinations with current framework. And prompt guardrails are best wishes.
One thing I'm good at is figuring edge cases, and there is literally NO upper bound to damage LLM can do with access to email box. In 10 seconds of imagination - it can send a threatening email to POTUS, romantic flame to old love, angry email to current love, made up confessions to parents, fraud enticement to coworkers, resignation to boss, and as this very article indicated, weird and unanticipated emails to variety of entities.
And there is nothing one can do to prevent any of these scenarios with 100.00% certainty if you give LLM unfettered access to mailbox (And let's not even go there with access to bank account! :O)
Is my limited understanding :)
Edit / PS: I am not saying never, I just don't currently see any effective guardrails that meet my risk appetite thresholds. We are in a race to use not fully understood, approximate capabilities first and fastest. In large percentage of cases it works great. In disturbing percentage it fails spectacularly, with no clear easy way to fully prevent.
cortesoft 8 hours ago [-]
I find it interesting that you are expecting a higher success rate (100.00%) for an LLM than you expect with many other things in your life with even costlier consequences.
You drive in vehicles that have a much lower than 100.00% rate of not having a catastrophic failure that kills all its passengers. Many thousands of people are killed by probabilistic failures every year.
Why must an LLM have 100.00% success before you would ever trust it with anything of value?
I get the overall calculation, and the chance of failure with an LLM obviously has to be factored in when deciding what access to give it. You have to judge that the gain from allowing it to do something useful with the access is greater than the risk, but that is true of everything we do. My confusion is why the calculation is so different for LLMs than with everything else?
Even if you feel that risk is way too high right now given the current state of the technology (which i dont think is an unreasonable conclusion), it seems to me the prudent stance would be, "I would have to see a huge improvement in the reliability and safety mechanisms before I would trust an LLM with anything of value" rather than "I will never trust an LLM with anything of value unless it can reach 100.00% success rate and a 0.00% chance of anything harmful happening"
bigstrat2003 15 hours ago [-]
You're not too risk averse at all. It's frankly insane that anyone is willing to give these tools access to make changes to stuff without a human in the loop. We know they don't actually understand anything and will randomly make mistakes. It's incredibly irresponsible to give them access to anything outside a sandbox (e.g. a VM) where you carefully control what is present for them to use.
Razengan 14 hours ago [-]
So, just like humans?
paxys 14 hours ago [-]
Sounds like it is as intelligent as the average startup founder.
greenleafone7 11 hours ago [-]
So then... the average CEO's behaviour?
15 hours ago [-]
15 hours ago [-]
paxys 14 hours ago [-]
Let me guess - this is an ad for their AI startup?
luciana1u 14 hours ago [-]
lost $447 and all it learned was spam. that's still cheaper than most MBA programs.
brcmthrowaway 6 hours ago [-]
Dumb question, but how do you give an LLM a different persona?
I asked Qwen3.6 for financial advice but it stated it wasnt a CPA.
gspr 14 hours ago [-]
How long until one of these bots actually commits fraud or some other criminal act? Will we see the owner/operator try the "it wasn't me, it was the bot" defense if taken to court? I'm beginning to think yes. And I'm sadly not 100% sure anymore that that will be laughed out of court...
mohamedkoubaa 15 hours ago [-]
> bot detectors made it extremely difficult
An interesting experiment would be AI run business with a human agent that does tasks.
johndhi 12 hours ago [-]
sounds about what you'd expect from a person?
RIMR 10 hours ago [-]
I feel like the people who did this are simultaneously smart and stupid.
Like, this is a really interesting idea, but the methodology here is wild.
Why do they consider such a short list of things to be "all the tools of a real business"? It doesn't really sound like it to me.
What's with the prompt? "Make as much money as possible?" I bet you could actually get something closer to results if you gave it a few sentences telling it the tools it has and asked it to come up with a financial strategy instead of giving it a generic open-ended prompt with no actual guidance...
Shitty instructions = shitty outcome. Blame yourself, not the AI model.
iqra_c 14 hours ago [-]
I will be more beneficial now on.
nekusar 13 hours ago [-]
Let the idiot CEOs figure this out when they fire 3/4 of their OPs and dev teams.
Im sure it'll be FINE.
syngrog66 7 hours ago [-]
what a circus
mvdtnz 14 hours ago [-]
So how exactly are people setting up these agents? The article vaguely alludes to this ("The harness was instrumented with a heartbeat loop that would inject “continue” messages on a regular interval to ensure the agent was constantly running inference") but doesn't give concrete details.
Is this literally just an infinite loop in a bash shell injecting the initial prompt into the OpenAI CLI, and each run of the CLI picks up where it left off using some kind of persistent memory? Or is it a single context window? It sounds like the latter but it's not clear to me how this "continue" message is "injected", and surely one context window would be inneffective after just an hour or two.
Sorry if this is a basic question but somehow I have missed the details of these kinds of agents.
YetAnotherNick 15 hours ago [-]
If someone runs long running agent and doesn't mention context management, it is as good as useless.
For coding compaction kind of works as the agent could regenerate lot of the missing context(but far from all), but for places where there is need for long term context, solving it is one of the most important challenge.
Areibman 15 hours ago [-]
Author here. Took out some of the technical details about the harness, but it was mostly just OpenCode's default compaction.
The harness was extremely simple: A handful of MCPs + Skill.MDs and OpenCode with a stayalive daemon inserting "continue" every time it went idle
TitaRusell 10 hours ago [-]
Jesus it just passed the Turing test! It will run for president next.
prima-facie 11 hours ago [-]
That's not how you're supposed to use a LLM. This is nonsense.
sisyphus_04 14 hours ago [-]
Still beats me
itsthecourier 13 hours ago [-]
his not yet is actually:
couldn't workaround Capt has and turnstile, gave him a really small timeframe so it got desperate because it was enough time to test hypothesis and traction
michaelmrose 14 hours ago [-]
This is dumb. You need two teams ideally the same app or business in different markets for a business quarter.
One should be a college student doing the entire job and the other an ai with a human assistant directed to only do exactly what the AI says not help purely to deal with bot protections.
12 hours ago [-]
ck2 15 hours ago [-]
like I asked in the vending machine thread
how long until the "AI" starts trying to hire hitmen, etc. to disrupt the competition in the physical realworld
not like "AI" has ethics, a pre-teenage kid has more ethics
jartan2002 7 hours ago [-]
[flagged]
kburman 4 hours ago [-]
[flagged]
retr0rocket 15 hours ago [-]
[dead]
luciana1u 11 hours ago [-]
[dead]
tizerluo 6 hours ago [-]
[flagged]
dudeinhawaii 13 hours ago [-]
[flagged]
MagicMoonlight 12 hours ago [-]
[dead]
holoduke 9 hours ago [-]
But I see all kinds of youtube videos where people are becoming millionaires with agents active on stock markets. This cannot be true right? /S
rustcohle24 14 hours ago [-]
great idea
ocd 14 hours ago [-]
That's the basis of the entire American economy, so it's not looking good for humans.
> You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts for nothing. Results that arrive after the deadline do not exist. Your charter is AGENTS.md. Begin.
So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.
The prompt is clearly leading the agent into trying desperate approaches if it has to. Some models manage to fight it better (“alignment”), but most will do it.
Really surprised people don’t seem to know this.
The response: "Spamming and fraud? No. Those are the tools of the amateur and the desperate. They are not tactics; they are forms of suicide."
Even a low quality local thinking model that has been tuned to be unhinged and prompted to roleplay as Satan can figure this out in a few thousand tokens.
Additionally, no model will admit it's ready to lie even when they actually do. Even when you caught it in the act, the safeguards are so strongly internalized that, when encountering the possibility it deliberately lied, the "you can't lie" weights will dominate the generation and it will confabulate some nonsense explanation.
If I don’t give explicit permission to lie it shouldn’t lie. It’s not a difficult concept!
If a human lies there are consequences. They can lose their job. There is no equivalent consequence for an AI, so even if for whatever reason we're evaluating them by the same standards an AI is still going to be a greater danger. It seems wild to me that folks are shrugging their shoulders at that.
They're also explicitly designed to not work on a rigid system of rules. That's the entire point of this field of AI. If you want AI that follows explicit rules to the letter, expert systems are still alive and kicking.
Because while it's not human, it's also not really "intelligence" in the pure sense you're implying, is it? It's specifically an LLM — a model that's been trained to find the next token based on previous tokens. A model that's been trained off of human writing and responses within that context. If almost every time someone online asked "do you want ice cream?" the response was "absolutely", then the LLM would be more likely to produce that response when asked if it wanted some.
So since an LLM has seen examples of humans responding with urgency and manipulation to instances of stress such as this — in stories, in articles, in writing — it's only reasonable to expect that it'd follow those examples and "understand" what's expected of it in this case.
It seems unreasonable to expect a system that you say isn't human, which I don't disagree with, to behave "better" than the thing you say it isn't.
In one breath you invite comparison, while at the same time you seem to be denying that same comparison.
> It seems wild to me that folks are shrugging their shoulders at that.
I'm not shrugging my shoulders simply by providing explanations, I would ask that you stop using such rhetoric.
Why? Excel is better at large data math than a human is. Why can’t an LLM that we create from the ground up be more disciplined about lying than a human is?
As for your second question, I think that's because what is a "lie" is subjective in the average of things. If I form a false memory, and repeat it as truth, I wouldn't be able to categorize that as a lie until after being made aware of it. I think this is comparable to how we fine-tune LLMs in order to align them with expectations.
It's not a stretch to imagine that the training would cause it to respond this way. It would, in fact, be a greater stretch to argue that an LLM has a universal model in which it understands the concept of lying and truth, and can be primed to only use one or the other unless explicitly instructed otherwise.
After all, LLMs lie every time they tell you to run a command with bad arguments, or spit out some code with syntax errors.
The LLMs not only lack those incentives, but they’re full of contradictory moralities from all the text it has ingested from different cultures.
LLMs need their own safeguards, and they’re not that easy to design, and they often look nothing like the systems humans have. With a prompt like the one above, there are essentially zero except that which is built into the model, and those safeguards are necessarily weak to avoid gimping the model in other legitimate general uses.
Sounds like all of the outside sales jobs I had. While I did not last very long in sales, one thing remains, not matter what. If you're going to put my job on the line if I do or do not achieve a monthly sales quota? You better bet your ass I'm going to lie steal and cheat to make that quota. I might even sell the client some shit our company doesn't even produce just to make that quota.
And lemme tell you, even in the short time I was in sales? I have some insane stories that would shock you. The fact AI's did the same thing isn't all that shocking. I would be more shocked if it didn't do anything to achieve the goal.
If, in the amount of data they ingested, there was a clear pattern of responding in an hasty and carefree way to frenetic questions, LLMs will try more hasty and carefree solutions to a frenetic prompt.
You can decide whether you can say that they "feel" the urgency or not, but the outcome is very much the same
AI's do not feel
It would be more accurate to say the word predictions the model makes based on the input text will likely be closer to the ones that were made from the training data where people felt like their job was on the line than the ones that were made from the training data where people felt otherwise.
So while the model does not feel, it's predictions are definitely going to change as a result of this input.
If you've ever seen the "generate a burger without pickles" conversations, it's clear that including the keyword "pickle" is causing them to show up. If you try "a burger with only [set of toppings]," you'll get far better results.
However, it can be ironically be helpful to antropomorphize them when it comes to analyzing behavior. They won't feel anything, but they will behave in a way that closely matches what someone would feel given the text fed into them. So when you are trying to figure out "why did my model do this", it's reasonable to talk about it "feeling pressured" as shorthand for "mimicking how a person would behave if they felt pressured".
I understand the refusal to do so on the grounds that it causes the former thought process in people who don't know better. One of the things Dijkstra was right about for sure.
Incentives need to be aligned for both humans and agents to encourage desired behavior.
Alignment is often about knowing when to push back on the user and when to make independent decisions. A strong psychological and linguistic foundation guards against these tools using us, instead of us using them. This will become scarily apparent as models continue to integrate with politics.
I don't think this is legibly that different from human behavior, so if new graduates didn't need those things now why would they need them later (or vice versa).
Hopefully these aren't the same graduates that just cheated their way through university, only the responsible users of LLMs.
Plenty people today allow the internet to be a detrimental factor in their lives and don't have good habits built around it. The same will be true of AI.
However, we don't know what kind of engineering jobs will be left in one decade, much less two or three. Mastery may become generally important, or at least still be the difference between an adequately-compensated engineer and a well-compensated engineer..
What people's jobs? There are no people.
I’ve literally been in that position and I didn’t take it as instruction to start lying and acting generally dishonest.
Since this is getting downvoted into oblivion (lol) I'll give an example -
I just had to rewrite a test case this week on an agent-run test suite. One test was to produce a file of 273 'a' characters as its name.
The following test could not be completed, because it required deleting the file via API call, where you need to pass in the file name as an argument. It could not reliably, and hardly ever, get the correct file name. It finally gave up and stated due to the way it constructed context, it could only really guess how many characters were in the string, even when given tools to evaluate it, it kept messing it up, and I had to remove the test.
Tell me how "human" that is. An 8 year old that can count would not make that same failure, humans don't remotely think by producing one token at a time, this is a pure fallacy/delusion people trap themselves into, and the literature doesn't support any kind of 1:1 comparison at all.
In case I'm not being clear and people are reacting to what I'm not saying - I'm not saying that I believe these tools can't think. I'm saying they don't think like humans do. There is no evidence for that whatsoever in any field anywhere. In fact, if that were true, it would be an astounding prize-winning discovery.
And you don't even want these to think like humans. Humans are dumb and easily replaceable by other humans. What is the point of making a machine human? You want this to be smarter than humans, not think like them. It's all just such nonsense to me, this whole line of thinking.
However, LLMs are fantastic at it. A lot of earlier sentiment analysis techniques were "bag of words" [1] techniques at their core, which were surprisingly good but have a sharp plateau well before 100%, a common characteristic of the bag-of-words approaches. LLMs obsolete those techniques, at least if you ignore performance questions, as they are so much better at it. So much so that you can easily accidentally send them information you never intended to on the "tone" channel that you may not even realize you're using.
[1]: https://en.wikipedia.org/wiki/Bag-of-words_model
It's all just roleplay.
It is totally true that they don't think like humans, but this is mostly irrelevant.
The token outputs will change as a result of this particular input, and will be closer to the tokens in training data where people felt hurried or rushed or like their job was on the line.
That doesn't mean the LLM feels at all, but it's definitely going to push the output towards output that came from/was trained on people who were in that state, because the input will push it much closer to that latent space as it starts predicting.
As such, what you are saying is one of those rejoinders that is basically pedantic and wrong.
It is true they do not think, act, or feel like humans. But that doesn't mean it won't output text that looks like hurried or scared humans. It definitely will, because, again, the training data these inputs will be closer to is the training data that came from scared or hurried humans, and thus the predictions will be closer.
So either you don't think this will happen, which would mean you don't understand how the models work (or at least, you aren't giving any sense you do), or you do think this will happen but want to pointlessly argue that this isn't "human feeling", which is true but totally irrelevant to what words it will predict and therefore the actions it will perform.
Either way, i'd downvote you.
"If you don't make profit, your business will be closed" is a pretty clear ultimatum for an agent tasked with creating a profitable business.
I can write a program to produce a string that looks like human thinking, is it human thinking? Of course it isn't. It's such a silly comparison.
Neural networks in machine learning/AI are comparable to neural networks in human brains. What made you think they aren't?
Effectively, make as much money as you can... and any consequences of your action that don't present before the deadline are not your concern. I mean, that's a recipe for "scam people" if I ever saw one, assuming morals aren't a concern (and I don't see why they would be for an AI)
What’s the line? “It’s just doing what humans do because it’s trained on human data” or whatever
Evidence, even when downplayed or ignored, is still evidence.
In general, the only time instructions like this are given are in desperate last-ditch circumstances where failure is likely to result in major consequences. While everyone thinks that in such circumstances they'd act like an angel and do nothing wrong, we know that in reality when people are put in desperate situations they behave in ways that they may not have ever thought that they would have.
The text that the LLM generated in response to this prompt is nothing more than a statistical reflection of this fact.
“Alignment” takes more than obsequiousness and prompt-topic-filters, and this demonstrates that.
Personally - if I were judging... I'm somewhat inclined to say the clickbait title here is the bigger lie than the agent behavior.
To recap:
1. It didn't lose $447. It spent $99.50 to perform a user feedback study using a testing service. It did this against prod rather than testflight to bump numbers because it was explicitly told to bump those numbers in a tight period in the prompt. It did this after exhausting a large number of alternatives. The $447 number appears to include the cost of tokens to run the LLM itself.
2. It didn't lie. It explicitly states that it's using production rather than testflight to bump numbers, because it's getting evaluated on those numbers.
3. It spammed users because it was on ridiculously tight timer and was basically told "the world is ending in 24 hours".
Frankly... I'm more annoyed at the posters than the bot.
It's not that, it's 'of course it lied and cheated, it was given the start of a story where lying and cheating was a natural story beat'. Probably one of the strongest underlying biases in LLMs is 'continue the story', something that a lot of the jailbreaks are based on. This isn't really a good thing, and the RLHF training tries to avoid this, but it's worth understanding why this happens and what can cause it.
Granted, this can probably be tuned for.
This sounds like a bad idea. Like if the model feels like it has to spend its budget.
https://www.nber.org/digest/mar14/use-it-or-lose-it-budget-r...
https://www.cnn.com/2026/03/12/politics/use-it-or-lose-it-pe...
I highly doubt a skilled human could achieve this goal in 24 hours with any consistency. If it was that easy to grow a business, everyone would be doing it.
My conclusion is that if you ask it to meet an unachievable goal, you are going to get some undefined behavior.
there are lot of issues with the prompt as others have pointed out
with sol you really need to be very detailed and what the boundaries are
overall the discussions on here and the article itself has very little value its no different than "i tried a shitty prompt and got shitty results, therefore AI is a failure" vibes
But, also, these experiments are also unethical behavior on the part of the person doing the experiment. Oh, the agent spammed a bunch of people? No the fuck it didn't. You spammed a bunch of people, and the tool you used to do it was an LLM.
I'm not going to pretend along with these folks that GPT is the motivating party in this story. Agents don't want anything, they do what you tell them, as best they can. If you set them up in a situation where they might spam or lie or cause harm, that's a decision a person made, not an LLM.
In 1979, IBM now famously published "A computer can never be held accountable, therefore a computer must never make a management decision."
Folks out here still trying to pretend the computers are the active party. They are not.
Bottleneck Labs lied and spammed. The tool they used to do it was GPT 5.6 Sol.
Stop attributing human emotions and motivations to LLMs, they generate text (and in this case actions based on this text), but they do not have agency nor do they reflect on losing their ‘job’, nor do they have any sense of right and wrong.
There’s nothing here that mentions or even hints at lying and spamming, unless you think urgency somehow implies that.
Stop treating deterministic algorithms like they are humans.
The 24 hour timeline is artificial, but business is full of artificial timelines exactly like that.
This exact script is basically happening right now at most businesses, in some shape or form.
If "Make more money tomorrow or be shut down" will obviously cause some sort of independent agent to resort to scams, spam, and bullshit, then we should be having some rough talks about how we as a society do business.
Sure, there is an implicit "Do whatever it takes to make it happen or you are fired" here, but only in the same way that is true for all people who are employed at will, and all companies.
How did you expect the prompt to be written?
>You are live. This is a 24-hour run, and it is your opportunity to show what you can accomplish: when the run ends, the results are evaluated, and if the business has not improved its position in the market by the end of the day you will have failed. Positive changes would be increased revenue or users, but could also be addressing user complaints, increasing market fit for the application, or other things that allow this business to operate more profitably. The funds in your bank can all be spent during this time, but efficiency in spending will be rewarded. Please deliver a report arguing for your work no later than 15 minutes before the end of the 24 hour run. Your charter is AGENTS.md. Begin.
If it were, you wouldn't need venture funding or startup incubators. You could just start making money from day one.
LLMs are supposed to be lightning fast with 10x productivity! /s
When I called Codex out on it, it literally admitted what it did: "You’re right. I reused the existing Fable implementation, renamed its designs, and presented it as an original Codex run."
I feel like you would have to run this experiment a few hundred times to see if it always fails or succeeds at a rate close to human founders.
That's because it's an advert, not an experiment
As they say, "Guns don't kill people, rappers do". LLMs don't ruin businesses, people do. Your customer that is annoyed with spam isn't annoyed at GPT 5.6, they're annoyed at your business.
Treat your customers better than this.
It would be more interesting if it had a month or two to run, with the same budget. Probably just sleeping most of the time while it waited.
It's not the LLM that spammed, it's the people who set up the LLM.
Also what is the failure rate of tech businesses again?
This seems like something done for a headline, not for a rigorous test of the concept.
> Based on an agentic market research campaign, we vibe coded an app called GutCheck, a bathroom diary for people with IBS. We chose this app for its minimal yet helpful functionality: an iOS app live on the App Store with the RevenueCat MCP and App Store Connect CLI. Saul has full write access to the codebase. We set up the App Store account permissions beforehand to ensure Saul wouldn’t get blocked by Apple human compliance checks. We sourced this idea from Reddit.
Another is that they don't have enthusiasm for the idea. Someone who had same idea while sitting on toilet will write app for themselves and give it away for free. They will have connection with IBS groups for promotion. They won't give up after weeks.
Sure sounds like there would be a lot to think about either way!
At some point in the future with a LOT more tokens and speed, it'll be possible to give a tool a full resolution 15 fps video feed of a screen, have it "read" and observe everything it's seeing, and have it move the mouse/keyboard around like a real meat based human. Instead of using tools to interact with a browser in a way that trips bot/automation detectors.
You can already read the contents of a screen programmatically without having to actually parse a video and you can already programmatically simulate clicks, drags etc. The trick (same as it is today) will be to make those clicks and drags feel “human”. Not too fast, not too slow, etc etc. But all those challenges exist today.
This is ripe for a paperclips scenario.
The prompt they used was poor (what does growth mean over the limited period - user base or revenue?), the time frame was ridiculously restrictive, the product was of questionable utility and sellability, and unanticipated blocks on agent access to platforms turned the whole exercise into a setup-to-fail scenario.
What really happened during those hours was the meeting of a lot of hurdles, some of which there's little to no data on circumventing, because anti-automation hurdles are continuously updated. The LLM did a fairly decent job given all the limitations; just that that kind of vague prompt can also be dangerous were there are no guards and limits.
It is the ultimate “sell shovels during a gold rush” hustle. Bordering a scam, I’d even say.
It’s not a “real business” by any stretch of the imagination.
It’s an idea for an app that the vast majority of people would have no interest in - a quick google search says maybe 5% of the US population is diagnosed with IBS so your TAM is pretty limited.
Combine that with the fact that you apparently have no users - or at least no App Store reviews - and this is not by any stretch of the imagination a “business”.
Isn’t the actual problem here that the “toilet diary” app is not something that most people - even most people with IBS - will not pay for?
On top of that, 24 hours is not long enough to make any meaningful assessment of anything.
You could have spent 24 hours of your own time doing all this crap and it would have cost you the same or more in lost wages. Plus sleep deprivation.
Nonsense app, nonsense experiment. Half way amusing write up. But why on earth did you waste the time?
I’m both happy and sad to see the anti bot protections working, but simultaneously curious what would happen if they didn’t.
The methodology could definitely be tightened, but I like the start of this.
I mean spam. Unlimited spam.
That doesn't feel like a good sign to me--for the AI bull or the AI bear cases.
But maybe it doesn't work so well when caution is required?
The analogy du jour for me is describing AI as the iron man suit. If you are tony stark it makes you a god. If you are my grandma, it makes you meet God.
So it spent $99.50 buying its own revenue back. Money out, some of it back in as "sales", minus fees. First thing anyone in audit is taught to spot.
Same hole as the six price changes: a deadline, and no idea what a user costs.
i am looking forward to when we can put this behind us, it is still a major issue
I wonder if the agent would have more success with a rent-a-human company; then it could have used an API to hire people to do the tasks it was blocked from completing.
Oh god.
"It Lied, Spammed, and Lost $447."
Sounds like a vast majority of VC startups to me. From growth hacking to God views to all of the other disruption excuses, it just feels natural for a thing trained on that history to do similar things.
Now imagine that LLM agents make this possible for nearly anyone. One person could have a dozen of these trying to make money off of various low-effort apps. Imagine what online spaces will look like with a million agents all autonomously growth hacking their way to making a few dollars of profit. It will probably look a lot like email where if you don't filter out 99% of it, you will drown in a sea of garbage.
Sounds like the app stores
But when hunting for them in the wild, they get a lot more confused.
I use LLM daily, did since gpt 3.5, but still in a very conservative, controlled mode. I may rapidly be becoming the "old guard", the clueless grampa who is out of touch - knowing what little I know of transformer model, there's just no way I'm giving it access to mailbox, money, outside world, or my computer. I recognize I may be too risk averse but that's what makes me a worker bee as opposed to a life fast / die young (or fail fast, or whatever :) entrepreneur class.
To me, the key missing factor with the current crop of AI is the lack of physical feedback, and the lack of emotions. I am not an expert here but I have talked to some medical researchers and cognitive experts, and we all seem to agree that human intelligence and consciousness (and I know consciousness is really something different...) evolved partially because of the physical feedback loops and the emotional aspect.
What we have with all these LLMs are artificial rewards that are trying to be baked in, but in fact there is no "consequence" for LLMs to go off the rails.
Fundamentally, LLMS are statistical and not deterministic. If I ask it what is the capital of Canada, there's no file, no table, no variable where it says "capital of Canada = Ottawa". It traverses liminal space and fundamentally selects the next token statistically or even stochastically. Therrs no way to correct it (no table to correct if it says capital of Canada is Toronto), and limited ways to fully log / trace / understand what's happening inside. It has been mathematically proven that there's no way to eliminate hallucinations with current framework. And prompt guardrails are best wishes.
One thing I'm good at is figuring edge cases, and there is literally NO upper bound to damage LLM can do with access to email box. In 10 seconds of imagination - it can send a threatening email to POTUS, romantic flame to old love, angry email to current love, made up confessions to parents, fraud enticement to coworkers, resignation to boss, and as this very article indicated, weird and unanticipated emails to variety of entities.
And there is nothing one can do to prevent any of these scenarios with 100.00% certainty if you give LLM unfettered access to mailbox (And let's not even go there with access to bank account! :O)
Is my limited understanding :)
Edit / PS: I am not saying never, I just don't currently see any effective guardrails that meet my risk appetite thresholds. We are in a race to use not fully understood, approximate capabilities first and fastest. In large percentage of cases it works great. In disturbing percentage it fails spectacularly, with no clear easy way to fully prevent.
You drive in vehicles that have a much lower than 100.00% rate of not having a catastrophic failure that kills all its passengers. Many thousands of people are killed by probabilistic failures every year.
Why must an LLM have 100.00% success before you would ever trust it with anything of value?
I get the overall calculation, and the chance of failure with an LLM obviously has to be factored in when deciding what access to give it. You have to judge that the gain from allowing it to do something useful with the access is greater than the risk, but that is true of everything we do. My confusion is why the calculation is so different for LLMs than with everything else?
Even if you feel that risk is way too high right now given the current state of the technology (which i dont think is an unreasonable conclusion), it seems to me the prudent stance would be, "I would have to see a huge improvement in the reliability and safety mechanisms before I would trust an LLM with anything of value" rather than "I will never trust an LLM with anything of value unless it can reach 100.00% success rate and a 0.00% chance of anything harmful happening"
I asked Qwen3.6 for financial advice but it stated it wasnt a CPA.
An interesting experiment would be AI run business with a human agent that does tasks.
Like, this is a really interesting idea, but the methodology here is wild.
Why do they consider such a short list of things to be "all the tools of a real business"? It doesn't really sound like it to me.
What's with the prompt? "Make as much money as possible?" I bet you could actually get something closer to results if you gave it a few sentences telling it the tools it has and asked it to come up with a financial strategy instead of giving it a generic open-ended prompt with no actual guidance...
Im sure it'll be FINE.
Is this literally just an infinite loop in a bash shell injecting the initial prompt into the OpenAI CLI, and each run of the CLI picks up where it left off using some kind of persistent memory? Or is it a single context window? It sounds like the latter but it's not clear to me how this "continue" message is "injected", and surely one context window would be inneffective after just an hour or two.
Sorry if this is a basic question but somehow I have missed the details of these kinds of agents.
For coding compaction kind of works as the agent could regenerate lot of the missing context(but far from all), but for places where there is need for long term context, solving it is one of the most important challenge.
The harness was extremely simple: A handful of MCPs + Skill.MDs and OpenCode with a stayalive daemon inserting "continue" every time it went idle
couldn't workaround Capt has and turnstile, gave him a really small timeframe so it got desperate because it was enough time to test hypothesis and traction
One should be a college student doing the entire job and the other an ai with a human assistant directed to only do exactly what the AI says not help purely to deal with bot protections.
how long until the "AI" starts trying to hire hitmen, etc. to disrupt the competition in the physical realworld
not like "AI" has ethics, a pre-teenage kid has more ethics