Welp, it's now blocking me from doing extraordinarily mundane tasks because of "safety". I've been an Opus fan for a long time, but this instantly made me cancel my subscription and move to OpenAI (which I also assume will screw me soon enough). Chinese models are almost there for my needs, and I can't wait to switch to them and never look back.
This has been my experience as well. Grok will end tasks almost immediately and claim "Done!". It's definitely the laziest and most "dishonest" of all the models. The others aren't perfect, but I can't use Grok for any serious coding task.
> This is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.
Local models are definitely not as productive as SOTA, sadly it's not close yet. I do think someday they will be "good enough" to use, but they aren't today. Even the SOTA models barely code well, with Opus 4.5 being the first, good coding model.
That being said, I think it's absolutely imperative that we keep pushing local model performance. We need to continue to advance technology there and ensure that the model labs don't do regulatory capture in the name of "safety" (or anything else).
With the latest codex (weekly quota burn) fiasco I tried open weight alternatives for the first time. And tyeah... open weight models cant compete with likes of astra yet. But, my hope is that by the time I get my Mac studio at end of november an open weight models would have closed the gap (which i think is realistic at the speed of progress). Now its true a better gpt version will also be available then but it also seems the gap is shrinking with time so theres that.
> And tyeah... open weight models cant compete with likes of astra yet
I think this is true, but also misses that a lot of us are just doing basic flask apps with a react front end. We don't need astra; Something sonnet 4.6 level locally is perfectly sufficient 95% of the time, and maybe 99% of the time.
This. People have convinced themselves that the absolute frontier is what is needed, anything below it is an unacceptable compromise, and we seem to be speaking different languages when it comes to discussing model capability.
It's like watching a discussion about cars available to take on a 100km road trip. A new car gets released that is on par with a Toyota Corolla but it is dismissed as completely useless for a 100km trip because it doesn't have the seat massagers and air ride suspension that the new Escalades have.
The reality is that something like Sonnet 4.6 is still amazingly capable for so many programming tasks, especially if you already have some reasonable level of experience to steer it in the right direction.
And if you think Sonnet 4.6 is still worthwhile, then it seems undeniable that something like Qwen 3.8-27B is also worthwhile.
Local models can be widely used as productive assets. Yes the infrastructure of SOTA API models is engineered specifically for you to be that utility, but the blanket statement that local isn't up to par is intensely short sighted. Billions of tokens per month on local pays for the hardware when compared to sota costs per month.
I believe they can currently be used productively for non-coding tasks (classification, light summary)... but they definitely are not even close to SOTA when it comes to software development.
Defining productivity is a use-case scenario, and a wildly generalized assumption for most people in this argument. Local infrastructure doesn't need to be sota for absolutely every single need for a dev lab, but it absolutely can be delivered with non-api frontier class models.
Just to be clear, I'm specifically talking about coding. I think local models can help with productivity today, just not coding.
I'm also a huge fan of local models and think it's absolutely imperative that they continue to advance so we can move off of the Anthropic/OpenAI hosted models. It's important to accurately asses where we are in that journey though.
I think the issue is generalization, if you were more specific about which local models aren’t good enough for which tasks compared to which frontier models in your experience, it’d be a lot more informative
Like the other commenter, I'm confused about the 'just not coding' conclusion. I'm using Qwen 27B on a 5090 at > 100tk/s with 150k context (which isn't enough admittedly), and DeepSeek v4 Flash with 1million context on a gb10/spark. Both of which are performing surface level, and deep needle precision infrastructure architecture. They code 24-7, stupendously.
It would be interesting to hear more about how you’re actually using them. Do you have sophisticated feedback loops around the models so they can verify their work and converge on good solutions? And how do you decide what to give the 5090 vs the Spark vs a frontier model?
Correctness matters much more than speed to me, but if I can get both, that’s obviously very interesting.
I so want this to be true, but for the kind of coding I do (not Flask apps), it's definitely not the case. Like I said, SOTA models just barely, barely work for me. My projects are usually 100k-1M lines of Rust or Go.
The SOTA models now work really well in my codebases, but that's only been since Opus 4.5/4.6-ish. Prior to that, and with current local models, they simply couldn't work holistically and would just thrash around. Now I feel as if SOTA are approaching my coding levels if not surpassing it. I still need to guide on architecture, but I can see that going away within the next year or so as well.
At least we know this wasn't written by an LLM (unless it was GPT-2).
The best advice I can give on moving your career forward is to not try and get a promotion at your current job. Apply for better positions at other companies while you still have your current role. You'll have maximum leverage and unlimited time to choose the best possible option. Trying to claw your way up the ladder in a single company is the least efficient (and more often than not hopeless) strategy.
The 'job hopping advice' is always so one-dimensonal as if salary maxing is the only metric that counts. You should always take everything into account
- work climate and coworkers
- work life balance
- legal rewards for staying (This might not apply to the US but in most of Europe you get progressively harder to fire the longer you stay at one company)
This. I hate interviewing. I hate having to "learn" new people and corporate processes. I like being able to rely on intuition I've built up over time, both about my colleagues and the products.
Bouncing from place to place could have earned me more money over time, but I valued the stability of sticking around for longer stints.
It is still good advice at early to mid stages of career. You can get into senior level much faster, and potentially learn much more, since you're less pigeonholed into current role, and you see more of how other organizations work.
Interviewing sucks. I know, if you currently have a job then you don't lose anything. Other than precious time and I already feel like I don't have enough free time, I don't want to invest the remaining interviewing.
I have been in a few companies and yes, sometimes getting a promotion is literally impossible. I remember one company that didn't give me a promotion until I threatened to leave.
But if you find the right company, then you shouldn't even need to ask for a promotion.
And lastly, I would hope you are working for a company that you care about. That makes it more complicated to find a replacement
I think you need to be strategic. Leverage your network, don't just go in cold. Smaller startups are also much easier to get a job at than a FAANG.
> And lastly, I would hope you are working for a company that you care about. That makes it more complicated to find a replacement
This I absolutely agree with. I quit my last job for ethical reasons and indeed, the same problem I faced is so pervasive in our industry that it limits the pool of places I'd work. That being said, I think the advice holds. Jumping around is a much easier way to move your career forward than staying at a single place.
This used to be true during the zero interest rate days when you could leapfrog up the chain but I think the times have changed.
By all means keep your eyes open outside of your org, but it's increasingly hard to even get a resume seen by a human on the other side of the hiring process.
If you have a good reputation on your current team, that places you above external candidates as you look to move up in your organization. Don't ignore that.
“Your job isn’t your job. Your job is to find your next career step.” Paraphrasing from memory from Scott Adams’ “How to fail at almost everything and still win big.”
I haven’t followed that advice, because I quite like my jobs, and getting deep into a problem space is one of life’s pleasures. But, I do think it’s sound advice all the same.
its both. i hopped mostly because i got to see how other places build things, different stack, different ideas about how to run a project, and you just dont get that staying at one company even a good one. internal promo really depends where you are. small startup, sure its easy to get promoted but youre basically doing the same job with a fancier title. big co you can move around and actually learn stuff but then you have to deal with the politics
I couldn't disagree more. The frontier labs are already working with the military to kill people. What you're going to get (even more than we already have) is a two tier system where those with weapons and a proven desire to use them will have the best AI and us regular citizens will have the hobbled AI. A complete reverse of what would keep us safe.
The top comment when I clicked it does not agree with the narrative:
@rkozakand 10 months ago
The so called "southwestern motif" is found wherever people weave. Our native rugs in the mountains of Ukraine use the exact same motif, only in different colors. It is a natural outgrowth of weaving techniques, which encourage triangles. We also have ancient pottery from what is called the Trypillian culture, which includes spirals and what you call 'zen tangle' motifs, which we call 'bezkonechnyk'. This pottery looks very similar to Southwest American pottery, even being executed in the same colors. I therefore doubt that they are as tied to specific plants as you think, although patterns repeat in nature, and maybe we have similar plants.
He actually talks about that in the video and compares and contrasts the weaving based patterns, nature based patterns and hallucinogenic patterns. The datura pattern being nature based.
More likely it's not the same protocol that served these different cultures those motifs, it's the same server sending the same data over different protocols.
It doesn't even make sense that geometric designs would be influenced by datura.
I have read every trip report on Erowid in relation to tropanes.
Art influenced by datura would look like Clive Barkers Hell Raiser. The trip reports are incredibly consistently dark because the person taking the substance is being poisoned. Not figuratively, literally.
The more interesting thing if you look into it is how ayahuasca use might be a modern phenomena. Maybe a few hundred years old.
Even more interesting is how there is a mindset that would not want this to be true and would never accept the evidence for it.
That's not what he says. The geometric designs are influenced by weaving. The patterns that look like datura, an ayahuasca... are influenced by nature, and some of the patterns in Mexico are influenced by actually hallucinating (he calls those zentangle patterns). The natural patterns aren't what you see when you take the drug, it's what the plant looks like. Imagine a box someone is storing marijuana in that has a marijuana leaf on it.
> Understanding that this is policy is crucial for comprehending how Israeli soldiers are expected to dutifully erase entire villages without questioning orders. By saying “200 NAZA” rather than “200 non-combattants killed,” the dehumanization process is handily condensed to four letters. And yet, despite being miles away from the drone attacks, the soldiers directing the strikes hear everything. The IDF listens to all Gazan phones, so they know where people are and whether children are home. They know who’s in an apartment just before the missiles hit, and they hear the screams after the explosions. They even coined a phrase — “Where’s Daddy?” — to describe listening in on a child’s phone, determining when their father will be home, and then sending a missile to kill them all. As one solider matter-of-factly says about a 12-year-old, “You hack her phone, and then you kill her.”
> The initial reaction, when some coding agent does a piece of work that would have taken you a week in an hour and does it well, is to be very disheartened by it.
I think it depends on the person. I've been waiting for AI since I first saw Wargames as a child. As soon as I saw that agents could code at or beyond my level, it felt like the entire world was unlocked. Now I can explore so many ideas that I was bottle-necked on with my own time to write the code before.
I would consider myself a technologist and it's a very exciting time to explore what technology can do. I loved writing code, but mostly because creating something that tapped the power of tech was my goal. Now I can achieve that goal 100x and it feels like a superpower.
Well what I'm worried about is that as AI gets better, everyone will be able to use AI to produce things. Just describe the product, or just point to an existing product, and that's it. In other words, AI could became better than every developer while being cheaper. So you won't be able to get paid for your work as a developer and will need to either retire or find another source of income.
I do think we need solutions like UBI, universal housing and universal healthcare. I'm fortunate enough to have the skills and mindset that I feel comfortable in the direction we're going. Just like with all technology, those who are excited and having fun with it are better positioned than those who fear or resist it.
Well this is what I was hinting at in my original comment. The main determinant of how people feel about AI is whether it threatens their source of income because at the end of the day, that's far more important than everything else.
Sure, you will be slightly better positioned if you're excited about this tech, but the impact will be typically marginal, IMHO.
Not sure about UBI. There will be enough jobs for the foreseeable future, AI only threatens the "sitting behind a computer" class of comfy jobs. But this is a tangential topic...
This is exactly how I feel, and I have been trying to think how to articulate it.
I love programming, and have loved it ever since I read the first page of my BASIC programming book 35 years ago and learned how to write a PRINT statement. I remember running into my parent's room as an 8 year old, telling them I knew how to make the computer print something. It was so exciting.
I have never lost that excitement, but my love is not just about programming. It is about technology in general. I love the idea of having smart machines, and being able to use them to do stuff, and the power you can have when you automate something. I love being able to come up with an idea about how a machine can do what I want, and then how that means it can do it over and over without tiring.
I loved technology in science fiction. My family watched Star Trek: The Next Generation from the first episode until the last. I loved Data, I loved the ships computer, I loved the holodeck. I dreamed about having those. I was so excited when I spent way too much money to buy the very first Oculus Rift dev kit, because it felt so much closer to the holodeck than anything I had experienced before.
I have felt that way about AI. I was astonished by the first ChatGPT I used. It was amazing! The first programming AI I tried was github's copilot, and I was already astonished. It was able to complete the function I needed! How did it know what I wanted to do?
I continue to be amazed at every advancement. I remember the first time I tried Claude, and I had it write me a connect 4 engine in Rust (writing connect 4 engines is one of my standard 'new language' routines), and it did it!
I have felt closer to living in Star Trek these last couple of years than ever before. Telling the computer what you want, seeing the response, and refining what you are asking for is how I dreamed of using a computer since I saw them do that on the Enterprise. And here I am, doing it!
I like computers and programming because I like technology, and I like seeing it advance into areas that seemed impossible before. I love the fact that the 'Tasks' XKCD that I loved to quote for so many years when non-programmers would talk about a programming idea (https://xkcd.com/1425/) has become completely obsolete; now anyone could spin up a bird identifying app in a few hours. What was 'virtually impossible' is now routine.
THAT is the thing I love, not writing code. I mean, I DO love writing code, but I love getting my machines to do what I want even more.
I can't wait until robotics become more mainstream. I can't wait to have a machine that can clean my bathrooms and do my laundry and cooke me dinner. I can already imagine the routines I will develop, and the new ways of combining those skills I will come up with when these things are invented. Or hell, maybe I will join one of the companies working on those things and move that field forward.
There is no end to what technology can do, and I am so excited to see what is next.
https://en.wikipedia.org/wiki/My_Lai_massacre#Court_martial
reply