I’m curious what you are using. The free versions of chatgpt have been like that for me, but even Gemini flash with extended thinking, also free for a while longer, is giving me pretty reliable results as long as there training data out there to derive an answer from.
Comment on Bill Gates: People saying AI will create more jobs don’t understand what’s coming
treadful@lemmy.zip 11 hours ago
I still can’t decide if these people are delusional or I am.
Every single time I use an LLM it fucking “lies” to me or otherwise completely fails at the task. The people talking like this seem to me like they’ve never actually used it, or haven’t actually vetted the accuracy (like most AI users).
Maybe I’m just not using the “good stuff”. Or I’m not imaginative enough to foresee a near future where these problems are actually corrected and it becomes trustworthy.
I’ve never been so torn by a technological prediction.
realitista@lemmus.org 11 hours ago
tyler@programming.dev 10 hours ago
Claude can one shot tasks until you get a larger system then it completely shits itself. These models are nothing more than autocomplete, and they can’t hold large systems in their heads. Anthropic literally tried to rewrite all of bun using Claude, they said they did it and yet it still hasn’t released six months later.
Not_mikey@lemmy.dbzer0.com 6 hours ago
until you get to a larger system then it completely shits itself
This has gotten a lot better for me by having a “send out scouts” skill that has a lower tier model agent search through the codebase before it starts to plan. Has handled my companies giant monolith pretty well and even can handle cross repo features as well.
Yet it still hasn’t released six months later
Claude code has been using the new rust bun for ~5 months now and has been working fine, and bun 1.4 that released in August is using the rust rewrite
realitista@lemmus.org 7 hours ago
Yeah I certainly wouldn’t advocate building a whole business around code it wrote. But for small personal tasks it hasn’t let me down. Building custom server applications, desktop applications, Firefox add-ons, upgrading my homeassistant 10 versions over a couple weeks without letting anything break. These sort of things it handles pretty easily.
treadful@lemmy.zip 11 hours ago
I’ve not yet fucked with Claude. I don’t want to pay for it, and I really don’t like the surveillance aspect of these centralized systems. Mostly I’m using Gemini, whatever DDG had in their search results, and local models I’ve been fiddling with (like Qwen3.8 right now).
All more or less garbage once I get into the details of anything on the edge of my expertise.
SnotFlickerman@lemmy.blahaj.zone 8 hours ago
I got a free year of perplexity.ai which gives me limited access to claude sonnet and yeah, its honestly pretty good.
RyanDownyJr@lemmy.world 11 hours ago
I’ve been using OpenCode with whateverthefuck free models they have listed on there and they all seem to do fine with agentic tasks like building me scripts or executables to make my work tasks easier.
I used Gemini at the start with “Frontier Knowledge” and it seemed to do worse than the ones listed on OpenCode, but maybe that’s because i could only do like three prompts a week since I refuse to pay into an AI.
end of the day, its just LLMs writing code for me, but I cannot see how this would be useful for a large scale code base, but also #NotAProgrammer.
realitista@lemmus.org 7 hours ago
Are you using Gemini in flash extended thinking? (Not flash light) . I haven’t had many hallucinations other than cases where the training data it needs just doesn’t exist (cases where I can’t find the answers by googling either)
ch00f@lemmy.world 11 hours ago
How do you verify that everything it tells you is correct?
riskable@programming.dev 10 hours ago
You check the links/references it gives you.
Gemini does a pretty good job of this because it doesn’t seem to have much built-in knowledge. Instead, it just searches the Internet on your behalf and returns summarized results with links to where it got that specific information.
I use it to search for scientific research all the time and the summaries often aren’t detailed enough so I actually click on those links. I’ve yet to encounter a situation where it fucked that up (invented links that don’t exist) but I have heard about it happening.
So far, the summaries have seemed to be pretty spot-on when it comes to biology papers 🤷
Psythik@lemmy.world 10 hours ago
By writing instructions to insist that it double verifies every (non obvious) claim with a minimum of two sources. I also told mine to always assume that the initial prompt is missing crucial context, and to ask as many follow-up questions as necessary until it has enough information to provide the answer to the question I’m really asking. (For speed and efficiency you can even make it give you multiple choice options to click on.) Because sometimes the problem isn’t with the LLM, but with the user asking the wrong questions.
Using those two instructions alone, I’ve encountered considerably fewer hallucinations, and when I’m still not certain, I can simply click on the sources linked next to every single claim the AI makes.
ch00f@lemmy.world 9 hours ago
Right, but that sounds like you implicitly trust the LLM and only verify statements when they seem off. So your intuition is the final arbiter if truth?
StabbingSky@lemmy.zip 10 hours ago
It definitely has an element of garbage in, garbage out.
Mikina@programming.dev 11 hours ago
I’ve been able to find a workflow where it’s mostly correct and can handle most of my gamedev related coding without making too many mistakes. I still have to actually read through the code and pay some attention to what it’s doing, and if it misses something and goes on a wild goose chase, it’s unusable and I have to start over (so someone who didn’t know what they are doing would be cooked), but whatever.
Sure, it does require a lot of looping adversarial reviews, and my average token cost is around 3000$ a month (we have unlimited budgets and a pretty accurate tracking, and also definitely cheaper than consumer prices per token with how large company it is), which is actually more than my monthly salary, but it’s just a job, for a company and on a product I don’t really care about, and I can 1) keep slacking in my job while doing my own coding stuff and projects, and keep seeing how absolutely unreasonable the prices are if you want to get at least semi-submitable results.
Is it worth it? Lol, no. The whole team is loosing codebase knowledge, we’re getting bottle-necked by pending PR reviews that are just stacking up and no one wants to do, so we’re not even more effective, the cost is absolutely absurd and in no way near sustainable.
And that’s while the whole industry is in the “Uber pricing” phase, so it will get a lot worse. But yeah, if you can burn 100-200$ per a simple implementation task, then it can have a pretty usable results. And that 3000$ a month does not include our CI review bot, that does additional rounds of multi-agent council reviews.
halfapage@lemmy.world 10 hours ago
you give me hope
fonix232@fedia.io 9 hours ago
Your experience is pretty unique then.
Yes, LLMs make mistakes, but even small, self-hosted ones are pretty efficient today if you prompt them well. They're not mind reading software so you need to be able to describe the task and HOW you want it done, not just barf in some basic instructions like "write me a copy of Facebook but better".
Fishnoodle@lemmy.world 9 hours ago
It sounds like you’re saying people still need to be smart enough to use them properly… Which will be a problem as people rely on them more and more, and in turn become more stupid.
treadful@lemmy.zip 9 hours ago
No amount of prompt “engineering” will help when they outright make shit up.
fonix232@fedia.io 7 hours ago
Yes it does. Just need to go beyond prompt. Add reinforcement loops, make it test the solution in a separate environ it can't screw up in, and have it not just INVENT things ("give me X"), but research the topic and base its solution on the rules created by the research.
This is what basically the Claude harness (not the local but the remote harness you can't see) adds to the LLM what makes it so powerful and useful. Replicate those processes and even a small 4B mode will be incredibly capable.
flicker@lemmy.dbzer0.com 9 hours ago
Like back when we had to teach people how to google shit. Tedious.
kewjo@lemmy.world 8 hours ago
from my point of view it feels like most people are willing to trade their cognitive function for being lazy, which is a boundary i never want to cross.
AI will produce a lot of code but most of it is pretty poor quality as most training data is going to be poor quality code, there’s just always going to be more bad code than good to begin with just due to how difficult quality code really is to produce. I’ll give a hint, good code is usually small and succinct.
since my work started pushing AI live site issues have increased dramatically. turns out the person using AI looks like they have a ton more productivity but in reality that just shifts to whoever is reviewing the code. and to those who will say its the developer’s responsibility to review the code, yeah no shit, but if you ever work corporate you realize most don’t care as long as their managers think they are productive and just blame others for being bottlenecks.
Overall it just enlightened me to how bad the average developer really is, but i guess obtaining mediocrity is the sacrifice to make in the name of “productivity”.
story@lemmy.zip 7 hours ago
what if it trained on openbsd and it’s ports
hcbxzz@lemmy.world 7 hours ago
I think a lot of people just have low standards.
dwemthy@lemmy.world 8 hours ago
No matter how much I carefully structure a prompt, define specific behaviors in skills, and tweak the agent md files it will still just go do something I don’t tell it to or not do something it’s got really specific instructions for. We have to write all code by LLM now at work and I’m trying to do my due diligence to review code before putting it up for PR. 9 times out of 10 when I tell it to show me a diff before committing it silently runs git diff in the background and prints “that’s the full diff”. That’s with some basic “here’s what I want when I ask for a diff” in the base context.
peopleproblems@lemmy.world 7 hours ago
Tbf I think its the context limit that makes things hard.
1m token context is so stupidly low for all of the input we consider and filter in real time.
redballooon@lemmy.world 9 hours ago
He doesn’t take about what is, he talks about what’s coming.
jj4211@lemmy.world 3 hours ago
Problem is that Gates isn’t really “in the loop” and doesn’t have especially valuable insight.
His position in tech was always a bit removed from the core technologist work, and now his exposure is a telephone game with people that are as distant from the tech as he was.
It is really going around with no shortage of commentators spewing out guesswork, but Gates is given more credibility by virtue of his role 30 years ago.
treadful@lemmy.zip 8 hours ago
That’s a fair point. I just don’t know if the future they see can become a reality.
redballooon@lemmy.world 8 hours ago
I said that a year ago, and half a year ago, too, expecting the S curve to hit and the technology to get to some upper limit. But instead we got the agent loop, and models that make really good use of it, the releases only get faster and faster, and real improvements with each one, either faster and cheaper, but just as good, or actually a good deal smarter.
Even should the S curve start to turn towards slowing down now ( and it doesn’t look like it), the ceiling that it’s going towards is so high, I am with Bill Gates here. We are not prepared for what is coming.
treadful@lemmy.zip 8 hours ago
I’m seeing signs of behavior that is more than just a good auto complete, and think there is actual intelligence.
I suggest you be real careful not to anthropomorphize these systems. To me this sounds almost like AI psychosis.
riskable@programming.dev 11 hours ago
Give us an example of some of the prompts you’re using and what LLMs. I’m curious if it’s a use case difference or you’re using the dollar store’s customer service AI to try to help you with your coding homework.
treadful@lemmy.zip 11 hours ago
Here’s a very common response to anyone that suggests they had a bad time with LLMs. You just aren’t using the right model. You didn’t ask the right questions. You didn’t give enough context in your prompt.
It’s not the fault of this infallible AI, it’s PEBKAC.
Nonsense.
OwOarchist@pawb.social 10 hours ago
Forgot to tell it ‘make no mistakes’.
StabbingSky@lemmy.zip 9 hours ago
Prompting an AI is very much a garbage in, garbage out type thing. Just like with any tool, you need to know how to use it properly to get the results you want.
Hell, some of the things I’ve seen people ask AI would confuse a human too.
Not_mikey@lemmy.dbzer0.com 6 hours ago
I mean yeah, it’s not magic, it’s a tool and it does take some skill / know how to use them correctly, and the people making those comments could be trying to teach those skills.
Like if someone said they had a bad time with Linux you’ll get similar questions and suggestions about their setup.
treadful@lemmy.zip 5 hours ago
I’ll concede that using LLMs in a useful way may take some skill. However, that’s not how these things are presented to everyone and it doesn’t reflect the reality of how the majority of people use them.
riskable@programming.dev 10 hours ago
Uh… I was just curious because sometimes it’s fun to see how the LLMs screw up. e.g. rocks on pizza.
It’s not the lack of evidence presented, it’s the person that asked the question that’s the problem.
treadful@lemmy.zip 10 hours ago
No you weren’t. You literally suggested I was using a “dollar store’s customer service AI to try to help [me] with [my] coding homework.”
flicker@lemmy.dbzer0.com 9 hours ago
I’m in college and for one of my classes we actually have a prompt we feed to ChatGPT so that ChatGPT creates an excel spreadsheet for us as a step in an assignment process. The whole spreadsheet.
The class is for healthcare and I won’t get more specific, but yes, you can use AI today to make whole ass Excel spreadsheets.
(One of the steps in that particular assignment involves fact-checking the spreadsheet, btw, but about 95% of the time all the answers are word-for-word in the book, 4% they’re taken from the web and have more up-to-date info than the book has, and 1% is blatant lies.)
aesthelete@lemmy.world 6 hours ago
The class is for healthcare and I won’t get more specific, but yes, you can use AI today to make whole ass Excel spreadsheets.
lol, it amazes me what some people are impressed by
jj4211@lemmy.world 3 hours ago
Yeah, so far when someone shows me a concrete example of this amazingly complicated thing that GenAI helped then realize, I think “really… that’s it?”
It’s a decent reminder that a lot of people have more modest needs and a lot of the enthusiasm can come from that, and maybe that’s more understandable.
teslasaur@lemmy.world 10 hours ago
I use AI all the time to parse log files for errors. Or write up simple scripts to do things that are one-offs or test of concept. Most, if not all succeed. I made a parser that translates the config of one brand of switches to another one, worked perfectly.
In what way are you using an LLM? It sounds like you’re asking it moral questions, to which it of course can’t give you an answer in any sort of objective sense.
jj4211@lemmy.world 3 hours ago
Generally I’ve found that:
If the facts are painfully obvious from a simple web search, then GenAI has a decent chance of getting it right. This can be useful if you can’t recall any “key” words well and the GenAI can craft several searches and get there.
However, if it does mess up the facts, the result looks superficially the same as “correct”. So while it can give accurate data, you always have to double check. This can still be useful, as finding the right search terms can be a decent help.
In coding, sometimes in some situations, you can have requirements that are absolutely testable, and thus you can have the models retry and retry until it works. This isn’t always feasible. Even when it seems feasible, you may screw up the criteria, or the GenAI when enough freedom disables a probablematic test rather than solve it, and it likely will generate code that’s not really fit to modify. There are a lot of situations where this is useful, but it is infuriating that non technical people and even some low skill technical people assume this is always the case.
Then when you get away from facts mattering, it gets “better”. Example, someone jokingly asked for one to “make gta6”. After a while it came back with a GTA 1 clone. Lots of people were impressed, because whatever it did could be considered a success even as it obviously didn’t match the expectation. The operators also like to GenAI some webcomic, where it is a fiction. They almost always didn’t have any interesting thought going in so they tend to be crap, but “correctness” didn’t matter.
Of course, also making fakes. Supreme case where looking correct matters but being factually correct does not matter at all. GenAI above all else “seems” correct.