Bringing data science back to AI - https://t.co/Zrmp6LRd9c About Me: https://t.co/P6WyeKkyTa
RT Sasha Sheng (Hiring) 🫶🏼 Lots of access are being granted from our discord - so if you are wanting access please check out there. https://discord.gg/typesafe
Just updated our AI evals FAQ with 5 new questions, 48 questions & answers total! New FAQs just added: - Do I need a reference answer or rubric before annotating data? - How can I do evals when traces contain sensitive data? - How do you review a trace that is really large? - How much context should I give a LLM judge? - What should I do when my “gold” eval dataset becomes stale? It's all here: https://hamel.dev/blog/posts/evals-faq/
RT Modal Runtime speaker lineup is live! We're bringing together experts covering AI infrastructure, applications of AI in science and robotics, the future of software engineering, and more.
🌶️Codex can do a anything Muse, Bot etc can do quite easily I get the appeal in the form factor “less is more” but if you are already a codex desktop and are a power user (computer use, remote access, thread management, voice control etc), YAGNI I still like playing with the other things for educational purposes but also need to avoid tool sprawl unless there is a real benefit
RT Quinn Slack Amp is now free to use when you bring your own compute and model subscriptions/keys. No more limits or fees for BYOK. https://ampcode.com/news/free-agent
It's cool that there is more interest in eval tools! Some opportunities for improvement: 1. Right now, this workflow tries to quiz you up front about your recollection of your experience with a plugin. It would be better if it was more "in-situ", meaning you could give feedback on the plugins as you are using them. 2. It goes off and builds datasets and judges automatically based on what you tell it in up front as well as what's documented in the plugin. I would like to see it try to do error discovery with you first to allow you to annotate real traces / session history etc so you can figure out what's not working better. 3. It was clunky to do this against a skill that wasn't a plugin. I had to ask claude to set it up for me and it took 10 minutes to figure that out. 4. I had to stare at the generated HTML for a while before I could understand what everything meant. There is some onboarding experience that is missing but I'm not sure what that is yet. I'm sure this will get better over time and the things above seem doable. It's also positive to see affordances for evals directly in our agents!
New in Claude Code: claude plugin eval See what value your plugin is adding, or if it needs more work. You can create test cases, run your plugin or skill against those test cases, score those runs, then run each case again without the plugin to see the differences.
RT vicki Corollary to this for software engineering is that I don’t know any good engineers who don’t tinker around with stuff that interests them in the field in their spare time
I can’t stress enough how little an idea matters compared to the agency of the people executing the idea. I have had the privilege of knowing and sometimes even working with some of the most successful people (by various metrics). The difference between mediocre and excellent
View quoted postRT Shreya Shankar Really stoked to soft launch what my new group is working on at SF Systems!!
SF Systems is back with a night of talks from researchers bringing their work to the real world! @sh_reya will share her latest work on domain-specific inference engines and @parker_ziegler will show us what programming can look like beyond text! https://luma.com/mzkxb97z
View quoted postRT Vik Paruchuri If you're blind, most PDFs are a garbled mess - and having one fixed can cost >$100. Our API now makes them usable for pennies.
RT Florian Brand in FrontierSWE, some tasks come with gold outputs so the model has something to work with (which are diff from the tests, so they don't matter tbf) and what do some models (like Grok) do with such things? thats right, hardcode these values so the local tests pass
RT Hugo Bowne-Anderson A lot of my working life over the past year has been preoccupied with harness engineering. When @HamelHusain introduced me to @doesdatmaksense to chat about how evals belong at the center of harness engineering, I immediately wanted her to write a guest post for the VG substack. Check it out to take your agent evals to the next level: https://hugobowne.substack.com/p/how-evals-are-central-to-harness Thanks, AS & HH!
RT Jay Alammar Over 300 original figures! An Illustrated Guide to AI Agents is now OUT on Kindle and other Ebook stores! Presses are rolling, print copies ship in a few weeks. @MaartenGr and I spent the last year and a half making the agent stack legible: memory, tools, planning, evaluation, multi-agent systems, and code agents. We're ecstatic to finally be able to share it with you! Agents are being adopted faster than they're being understood. Our bet is that a small number of concepts will outlast everything else in this field, so we spent the time identifying them and explaining them properly. Concepts and code. You build a working agent yourself along the way. Order here: Amazon: https://www.amazon.com/Illustrated-Guide-AI-Agents-Concepts-ebook/dp/B0H381PNPR/ O'Reilly: https://www.oreilly.com/library/view/an-illustrated-guide/9798341662681/
RT Gergely Orosz What’s the problem w using LLMs to write stuff you want other people to read -emails, articles, posts- and why are people (like me!) so incredibly hostile that they will block you/mute you/block your domain, to never read anything from you again? Bryan spells it out. Read it:
The revolt of the reader https://bcantrill.dtrace.org/2026/09/05/the-revolt-of-the-reader/
View quoted postRT Hugo Bowne-Anderson Welcome to the revenge of the data scientist! People spent the past few years declaring data science dead. Then we started building products around agents that generate more data, noisier signals and plausible-looking outputs nobody knows whether to trust. Someone has to reason about all of that, debug it and work out whether the system is actually working. And guess what... that’s data science! I’ve taken @HamelHusain and @sh_reya's AI Evals for Engineers & PMs course several times now. They’ve rebuilt the whole thing from scratch for the age of agents, so apparently I’m going back again 🤣 The new cohort starts today. Hamel and Shreya have very kindly given the Vanishing Gradients community 25% off: https://vanishinggradients.short.gy/hamel-evals Come learn with me!
. @isaac_flath and I are going to see if the slop prompt is good and do some evals. (also first time trying to stream haha). https://x.com/i/broadcasts/1DGLdZNYNNyGm?s=20
I love how Teresa is merging product discovery with evals. I think its a powerful combination! I started learning about product discovery b/c of Teresa and I think its a great skill for any engineer as it helps you focus on what to build. Highly recommend checking out her work.
AI evals have been the "it" skill for product teams for over a year. I've even called evals a new discovery habit. But I still meet product teams who only have a vague idea of what evals are. And it's not their fault. Most of the writing on this topic is intended for engineers
View quoted postRT Ivan Leo thought it was funny i asked @bot to teach me more about post training and it just created fake sft samples and told me to look at the data and do data labelling can't run away from @HamelHusain 's just look at your data haha
RT Cognition GPT-6 Astra is coming to Devin. On FrontierCode 1.1, Astra performs within 0.4 points of Fable 5 at a 64% lower cost. It also sets a new SOTA on our internal testing benchmark, generating more comprehensive tests, clearer reports, and better video evidence.
RT Hamza Tahir Love @aiDotEngineer but can't watch 1000s of videos on YouTube? Me neither. My guy @strickvl built a website that indexes all @aiDotEngineer talks and makes them easier to consume. There are now 1,100+ talks in the archive. Each one has a short TL;DR, a proper summary, timestamped ideas that jump into the video, quotes, references, and tags. It is how you can decide whether to watch a 45-minute talk now, save it for later, or just take the useful bits. Interested in @dexhorthy's bangers? Go here: https://aietalks.com/speakers/dex-horthy. Maybe @HamelHusain and evals are interesting: https://aietalks.com/speakers/hamel-husain Or just go to a pack of talks like coding factories evals etc: https://aietalks.com/packs/coding-agents-on-real-codebases It'll be live updated, so benchmark it in time for Paris and NYC. @swyx hope the community likes this one! Bookmark here: https://aietalks.com/
RT Matt Pocock /show-me is a phenomenal skill Makes PR descriptions extremely easy to read Basically a toolbox of "nice ways to look at code" Nice work, @dexhorthy https://github.com/humanlayer/skills/blob/main/plugins/show-me/skills/show-me/SKILL.md
Amazing advertising
GOPHRS is here. AI gateway migration is mandatory. The arch is final (as in "no more feature requests" final). Please watch the full architecture briefing before submitting questions.
View quoted postRT Hugo Bowne-Anderson "Everyone is building this AI agent. DON'T" -- @HamelHusain full conversation here: https://www.youtube.com/live/QCBLUokyvHA?si=IcEUMrgWF7pfKyp-
The suggested prompt to remove slop is itself, pure slop If it works fine, but feels funny to me ngl
I guess we are back to magic phrases again for prompting Re: “mannered prose”
RT Hugo Bowne-Anderson An AI data agent tells you net revenue was $4.21M. It doesn't show the metric definition, source query, assumptions, or what it could not verify. The only way to trust the answer is to redo the analysis yourself. That is a product problem. I’m going live in 10 minutes with @HamelHusain to work through what AI products should expose before users can trust their outputs: provenance, visible calculations, trusted starting points, diffs, contradictions, and smaller units people can inspect, edit, accept, or reject. Hamel has spent the past three years focused on AI evals. We’ll dig into why “hard to eval” is often a product smell, and how designing for verification can produce stronger eval data. Watch live: https://www.youtube.com/watch?v=QCBLUokyvHA
RT Lenny Rachitsky I'd always thought AI was terrible at design, but after reading today's 🤯 post by @anshuc, I realized I was just doing it wrong. "AI models are capable of amazing creativity, but that creativity gets stifled. LLMs are trained to be next-token predictors: they look at a sequence of text and predict what typically comes next. Great design is exactly the opposite of this. Great design bends the rules and delights users with memorable, unexpected choices." @anshuc led design and engineering teams at Apple for 12 years. In his words: "Most people only see 1% of AI's creative potential. I want to show you how to tap into the other 99%." His 8 techniques for breaking out of the 1%: 1. Use seed strings to inject variety 2. Be much more ambitious with your prompts 3. Create positive feedback loops with subagents 4. Use image generation to enrich designs 5. Use video generation 6. Cut out elements that don’t add value 7. Remove AI tells 8. Rewrite copy by hand Read the post here: https://www.lennysnewsletter.com/p/how-to-turn-your-ai-into-a-world P.S. This design was made by AI 👇
RT Hugo Bowne-Anderson chatting with @HamelHusain tomorrow about * why "it's hard to eval" is a product smell * how agents have changed the eval landscape and how @sh_reya and he have updated their course * whether data science is dead in the age of AI agents or not Register to join live or get the recording afterwards: https://luma.com/7lng145m
Curious why machine learning is not applied to the X timeline more? If somebody is replying once a minute for 8 hours they should be flagged. When I look at reply bot replies its super obvious
Question for people who run autoreply bots here: are you not worried about the negative impact they have on your professional reputation? Anyone checking your profile here - a potential future employer for example - will instantly be able to tell you automated replies with a bot
View quoted postRT Simon Willison Question for people who run autoreply bots here: are you not worried about the negative impact they have on your professional reputation? Anyone checking your profile here - a potential future employer for example - will instantly be able to tell you automated replies with a bot
Teaching is hard but worth it when you see this 🥰 @sh_reya and I are refreshing the material once again to incorporate the latest eval techniques in our next cohort https://evals.info
Activity on repository
hamelsmu forked hamelsmu/reasoning-from-scratch from rasbt/reasoning-from-scratch
View on GitHubDon't fall for this rage bait 😅 DS gives you the tools and judgement to make sense of noisy signals and stochastic outputs For AI -> lets you measure if an AI product is working even though the outputs are text. You have to design experiments and test hypothesis that account for noisy signals. More on this here https://hamel.dev/blog/posts/revenge/
Data science feels so dead now that it's like it never existed at all. AI seems to have killed it
View quoted postRT Bryan Bischof fka Dr. Donut This may be the most misaligned with Kareem I’ve ever been. Data science is the best it’s ever been. I’ve been doing events and podcast episodes and production work in ai+ds for the last >3 years and honestly it’s a helluva time. It’s not *the same* but it’s a great time. As I start teaching again next week to DS masters students, the question is not “is this worth teaching” it’s “how do I find time in class for all the new opportunities!”
Data science feels so dead now that it's like it never existed at all. AI seems to have killed it
View quoted postRT rahul are we just supposed to forget what they did to windsurf?
Cursor has been a trusted partner of Anthropic since Sonnet 3.5. We’ll continue to increase compute to support Claude models in Cursor and are excited for what comes next with them at SpaceX.
View quoted postRT Sarah Catanzaro Eval developers are the next analytics engineers; people who will set standards that ultimately govern how models are trained and deployed. I expect to see many more companies hire for this role.
WE NOW HAVE A HEAD OF EVALS the work he's doing is honestly gamechanging. cannot wait to show it to you
View quoted postReally great use of tokens. Highly recommend. This is an AGI level task b/c have to click forms and such. Codex computer use flying through it
RT jacky we're hiring cracked evals/benchmarking ppl research-oriented role super tough hairy unsolved problems those that thrive in unknowns preferred pls dm me with your actual work (not resume), will fast track you
100% of the replies to this are AI bots Super disappointing
Really interesting new blog post from @openai for several reasons: 1) Shows an example of building with WebMCP, meant for when you want agents and and humans to collaborate on using a UI (like co-editing notebook cells). It's different than MCPs or APIs in that its exposed
Really interesting new blog post from @openai for several reasons: 1) Shows an example of building with WebMCP, meant for when you want agents and and humans to collaborate on using a UI (like co-editing notebook cells). It's different than MCPs or APIs in that its exposed directly through the browser. Read the post for discussion of the tradeoffs. 2) They created a new kind of notebook which works with WebMCP that prioritizes meeting people where they are: you bring your own coding agent and files are just markdown. The author uses it to curate runbooks or high quality examples of how to run foundation model evals on their infrastructure. Notebooks are good for this since they require tinkering with state of long running jobs interactively while taking notes inline. And its open source ✨ Blog: https://learn.chatgpt.com/blog/automating-repetitive-work-at-openai-with-codex P.S. this post is authored by Jeremy Lewi who isn't on X but here is his website https://lewi.us/
RT dex You can’t claim “the models are good enough that I don’t have to read the code”. Because if you’re not reading the code, then you can’t possibly know how much slop is getting in. We go live to @0xblacklight and @vaibcode for the scoop
RT Peter Yang I don’t promote other people’s courses often, but I want to recommend Hamel and Shreya’s top rated AI evals course. Here’s why I think it’s worth taking: 1. A 4.7-star rating across 900 reviews, with students from OpenAI and Google, is almost unheard of on Maven. 2. They’ve completely revamped the course and added a 24/7 AI evals assistant to guide you through the process. 3. They’ve personally helped me improve my evals, and their advice is consistently practical, specific, and grounded in real experience. You can watch my free podcast episode with them to learn more about their eval process before deciding: https://youtu.be/bdMHQLvtVaQ The next cohort of their course starts September 5 and you can get 25% off with this link: https://maven.com/parlance-labs/evals?utm_campaign=peteryang&utm_medium=affiliate&utm_source=maven&promoCode=PETERYANG
RT Florian Brand tired: hill climbing an eval by doing a synth env wired: hill climbing an eval by fixing the eval
RT Peter Yang There are two types of AI evals - tops down and bottoms up. From @sh_reya: “Think of top-down as: If you’re in a vacuum, just given the task description, what would you come up with? Claude does a very good job of helping you with these top-down evals. Bottom-up evals are the other half. When you look at lots and lots of sample outputs, what is your gut feedback that you want externalized into evals? “Claude is very, very bad at coming up with bottom-up evals. That’s all you.” 📌 Watch the full episode here: https://www.youtube.com/watch?v=bdMHQLvtVaQ
“The fundamentals for [AI evals] still apply. Start by looking at real data. What has changed is getting agents to help you look at it in a thoughtful way.” Here’s my new episode with @sh_reya and @HamelHusain, who have taught AI evals to 4,500+ engineers and PMs. I asked them
View quoted postI was wondering why Google is in its own special toilet tier in this video 🤣 Summary: it’s expensive as hell compared to everything else
😅Now I **have to** try it. Firing up the old model training rig that’s been collecting dust
@HamelHusain This is how the best adventures start. Just ask me and @tobi 😂. Many such stories! Never once have I regretted it.
View quoted postSneak peek of segments of the new evals course ✨😃
we've been iterating a lot on how DocWriter internally represents a user's writing style! the naive approach is to dump all a user's prior writing in context and pray the AI figures out how to sound like them. this doesn't really work. and the user can't just manually specify
RT Peter Yang “The fundamentals for [AI evals] still apply. Start by looking at real data. What has changed is getting agents to help you look at it in a thoughtful way.” Here’s my new episode with @sh_reya and @HamelHusain, who have taught AI evals to 4,500+ engineers and PMs. I asked them to audit the evals I built for my creator skills live. They then demoed a free skill that you can use in Claude Code or Codex to build reusable evals from your feedback. Some quotes from Shreya and Hamel: “Bottom-up evals come from looking at lots of sample outputs and turning that into eval criteria. AI is very bad at coming up with them. That’s all you.” “The agent’s job is not to invent new feedback. But it can help you group and distill the feedback into actionable rubric criteria.” “All your competitors can point Claude at their product and say, ‘Find all the errors.’ What matters is how much taste you can infuse beyond that.” 📌 Watch now: https://www.youtube.com/watch?v=bdMHQLvtVaQ Thanks to our sponsors: @WisprFlow: 4x faster than typing with your voice https://ref.wisprflow.ai/peteryang @linear: The AI agent platform for modern teams https://linear.app/behind-the-craft
It's incredibly hard not to get nerdsniped by Omarchy Looks amazing
RT Peter Yang In my next episode, I asked @sh_reya and @HamelHusain (taught evals to 4,500+ engineers and PMs) to roast the AI evals that I built for my creator skills live. They also showed me how anyone can use their free Error Discovery skill in Claude Code or Codex to turn real AI failures into reusable evals. 📌 Subscribe to get the full episode tomorrow: https://www.youtube.com/@PeterYangYT?sub_confirmation=1
WTF is "deep tech" its a term I only hear in investor speak and makes no sense to me
RT Peter Yang I'm on my way back to Vancouver to be with my mom but wanted to take a moment to celebrate crossing 100K subs on YouTube. Excited to share a lot more practical interviews soon to answer your most burning AI questions: 1. @HamelHusain and @sh_reya (AI evals experts) on how today's AI models have changed evals completely 2. @amoljain_ (Head of Product Engineering at Replit) on which vibe-coded apps have actually become real businesses 3. @ebloch (Product Lead at OpenAI) on how ChatGPT Finance can help you save both time and money 4. @poteto and Roman (SpaceXAI) on how the Grok Bot team uses Grok @bot 📌 Subscribe to get the episodes soon: https://www.youtube.com/@PeterYangYT?sub_confirmation=1
RT sarah guo .@gabepereyra, the research team at @harvey, and their partners are giving everyone building specialized intelligence a clear blueprint for the huge performance and efficiency gains possible from post training and in-domain data and workflow understanding own your intelligence!
Vendor posts that insult a specific competitor make me seriously consider signing up for the competitor Because they must be pretty damn important to name. Very rookie comms and marketing mistake
RT knut sat down and actually read @rosmine's research paper on this. https://deftwriting.com/research/distribution-fine-tuning first of all, i was guilty of having a "take" based on this post, pointing out how the copy in the screenshot is not a great example of what "good writing" is, despite scoring 100 on "human written" in @pangram. I recommend actually reading the work before commenting on it - because sloppy takes are as lazy as sloppy texts. (shame on me for joining the band wagon) but sleeping on it, i am grateful that @rosmine took the time to dive into this stuff. AI-slop fatigue is real, and we should support efforts to make agents produce communication that is clear and lucid, and doesn't feel overly synthetic but there are some assumptions in this, however, that is interesting to question when it comes to "what makes for good writing" as far as i understand, Deft is trained to have more diversity/variation in tokens, so less repetition of what we are recognizing as "AI-tells" (certain word and stylistic choices). And i totally agree with @rosmine that: "LLMs are not the cause of slop. Lack of effort/care is. If you spend days researching and planning a blog post, and put all the information into a detailed, well-structured outline, and ask ChatGPT to generate the post based on the outline, then the output will be interesting to read, even if the text has a lot of em-dashes." I argued the same in our eng blog announcement post yesterday: https://www.sanity.io/engineering/announcing-the-sanity-engineering-blog BUT! I still feel that this report (at least somewhat), but especially the various takes on it, conflates something sounding "human" with it being "good." Tricking @pangram doesn't make a text well written. Some reflections: - Making writing sound more "human" by means of adding more variation in word/style choices, doesn't make it better - What makes for a "good" text is highly contextual. If you are writing a recipe or instruction...
Announcing Deft, a new AI lab for better writing, cofounded with @jmrphy See the picture for launch announcement the Deft model wrote for itself Currently, 86% of user queries are fully human according to pangram. This is still a small beta model and it might make mistakes. We
RT Jonathan Whitaker I've officially left https://answer.ai I've had a good rest, with plenty of time for travel & tinkering, and now I'm thinking about what to do next. I've got a few ideas to share soon, but I'm also open to suggestions - feel free to reach out :)
RT ¯\_(ツ)_/¯ Re @HamelHusain @pangram @rosmine the iron law of tech twitter is when your employees start douche posting you are hitting some revenue wall and they need an outlet for frustration. pathetic stuff.
RT Alex Strick van Linschoten Re Good question. The main secret sauce is... you :) Basically we believe that evals are all about pairing domain experts with systems that allow them to encode their taste and judgement and put this all together in a workflow where you're improving your agent by seeing what's going wrong, using deterministic (where possible) evaluators to capture those failure points, and then using the replay etc to make sure that you've actually fixed things at the root. We're not really at the point where you can just automate evals fully without humans being involved (see @HamelHusain's recent post https://parlance-labs.com/blog/posts/auto-evals/index.html on some of the ways that can go wrong), but for sure tools (like coding agents, or like Kitaru) can help make this process as painless as possible.
New meme template thx to @BEBischof
A few months ago,@sh_reya and I released eval skills plugin. We iterated on it a bunch and recently made it better The biggest change is a new error-discovery skill. Give your coding agent a file of AI outputs or traces, and it builds a custom review app w/intelligent sampling. As you annotate the sample, the agent groups your notes into failure modes and finds related examples. We also added a start skill, which looks at your situation and routes your agent to the right workflow. It can help you find errors in a set of traces or audit an eval pipeline you already have. Writeup: https://hamel.dev/blog/posts/evals-skills/ GitHub: https://github.com/ai-evals-course/evals-skills
Hire John. I worked with him personally at Airbnb and he’s top 1%
I'm looking for my next role. I'm an AI engineer / PM, currently based in Berlin and open to relocating. Most recently Head of AI & Product at an AI fintech until the startup shut down in June. Before that: data science at Trumid and Airbnb, and Head of Data at Circ. Over the
RT John Enevoldsen I'm looking for my next role. I'm an AI engineer / PM, currently based in Berlin and open to relocating. Most recently Head of AI & Product at an AI fintech until the startup shut down in June. Before that: data science at Trumid and Airbnb, and Head of Data at Circ. Over the past couple of months, I’ve also been getting much closer to the model side: training three LLMs from scratch in PyTorch, post-training them with SFT and GRPO, then building the inference and serving stack, including KV caching, streaming and FastAPI. Everything is open source: models, weights, code, write-ups, and a playground with all nine checkpoints: https://huggingface.co/spaces/JohnEnev/modern-llm-playground I’m looking for an AI engineering or AI PM role, ideally somewhere I can stay close to both the model work and the product. Berlin, remote, or open to relocating. If you know of something, or someone worth speaking to, I’d really appreciate a pointer or repost.
Activity on hamelsmu/evals-skills
hamelsmu opened a pull request in evals-skills
View on GitHubRe: watermark - all that’s gonna happen is there will be tools and APIs to strip the watermark out In the end this will result in more friction instead of visibility
RT Shreya Shankar Interesting to see all the agreement. Check out the DocWriter project's plain-writing skill: https://github.com/docwriter-org/plain-writing-skill We recently added evals for the skill --- comparing a gpt-5.5 agent without the skill to with the skill (using an LLM judge to evaluate each criterion)
New term coined: ✨ Fablish ✨
@HamelHusain I also cannot understand Fablish. I feel crazy because everyone else loves Fable so I’m wondering what I’m doing so wrong
View quoted postI cannot use Opus 5.0, yes it can code but it can no longer explain what it did intelligibly It feels like its opaque internal reasoning dialect is now being used to talk to humans?
RT Shreya Shankar opus 4.6 feels like the last model that spoke english
RT Shreya Shankar Lots of new material in our AI evals course this fall. We put together a revised version of the 200-page course reader. Some of my favorite additions include: how to take an agentic approach to building evals, and how to think about safety, risk, and privacy as someone deploying a bespoke agent for their organization. Next cohort starts September 5th. Check out the full syllabus here: https://maven.com/parlance-labs/evals?promoCode=evals-info-url
Grok 4.6 just ranked #1 on CursorBench 3.2 Outperforming Claude Fable 5, Opus 5 and GPT-5.6 Sol on real-world coding performance And what makes this even crazier is the efficiency....the chart gives CursorBench performance against average cost per task, and Grok 4.6 is sitting
RT Isaac Flath Amp orbs are actually pretty useful. Here's an example. I started a thread in the cli working on my blog. Had it use a portal to host the site to show me the changes. I'm working on a technical post and wanted it to render jupyter notebooks nicer so it did that and I saw it work in the site portal. I moved to to the web UI because it was just kinda nice to have my website preview and coding agent both in chrome in the same app. I went back to work on my post, so I told it to make another portal for jupyter. This was nice because I don't have jupyter on this machine and I don't really like telling agents to just go install stuff on my local computer, but in an orb sandbox it's fine. It's on same machine so save of notebook reloads the preview site. And it all just worked. I could have done all this on localhost like I usually do. But it was nice because it was just a bit less friction than normal
Grok Bot is cool because it is "self-driving" (like Codex desktop). It can create new threads/tasks and communicate b/w them It's also a simple interface and decent computer use + nice iOS app. I found computer use lagging behind Codex slightly but its worth trying it, especially if you have a Cursor Ultra sub https://x.ai/bot
The model "just needs encouragement" is a UX issue. It's not something to be proud of
“overcast” describes cloud cover precisely, while “grey” might describe the light, sky, mood,etc. Yes, I care, and it matters.
There is no blog post that will win over developers re: watermarking AI
RT Nick Dobos Anthropic’s response to watermarks confirmed my worst fears. “Our watermarking method doesn’t have any practical impact on the quality or content” That is a lie. They literally say they change the wording in the first paragraph Anyone with an ounce of common sense about how psychology, writing, and legal documents works in the real world, knows this completely changes the meaning and content. This degrades and changes petulance. Look at the example they picked, it’s garbage. “Overcast” vs “grey” are completely different descriptions. That’s their best example, and it changes the sentence subtly yet dramatically. This is unbelievably dangerous This will kill people I am not exaggerating Anthropic just handed the keys to mind control the government They just gave politicians a fully legal back door that can be exploited to influence claude’s thinking to mind control millions of people without anyone ever knowing This is not okay.
We’ve written an FAQ to answer some of the questions we've received about watermarking. In summary: • We’re implementing watermarking to comply with the EU AI Act. Other major model developers have signed the same Code of Practice and will also be implementing watermarking;
View quoted post💯 “You didn't read the thing when you generated it, I won't read it when I'm reading it.”
new post: write for people https://vickiboykis.com/2026/08/12/write-for-people/
View quoted postRT Shreya Shankar It was great to present our data agent benchmark in the summer of evals series! Slides courtesy of co first author @ruiyingm1120, a second year PhD student at UC Berkeley!!
New session w/@sh_reya where she goes over a useful new eval called the Data Agent Benchmark (DAB). Data agents answer business questions like “which cohort had the highest churn?” that a data analyst would normally answer. DAB recreates the mess of a real data warehouse.
View quoted postRT vicki new post: write for people https://vickiboykis.com/2026/08/12/write-for-people/
RT Lambda Most teams lose sight of what exists outside the world of frontier APIs when it comes to their everyday work. Lambda’s @TheZachMueller sat down with @HamelHusain on when an open model is the right call, and how to serve it well once you’ve committed. It also builds on @_xjdr’s point that open weights can handle ~90% of tasks for ~90% of people. The rest stays on frontier work. http://youtu.be/Pg-IW5puuv0
Over the summer, @sh_reya and I hosted 13 sessions on AI Engineering topics like retrieval, post-training, inference, and evals. I've summarized all the sessions, organized by theme, with links to the source materials. Warning: I've tried to pull the most important ideas from each talk, so some notes are short (9.5 hours of sessions comes out to about 20 minutes of reading). Enjoy! https://hamel.dev/notes/llm/ai-product-engineering/
RT Alexis Gallagher Developing a robot by chatting with a robot https://x.com/i/broadcasts/1qKDzWzBqmkJV
RT Hugo Bowne-Anderson Fuck your skills. @HamelHusain is the guest I’ve had on Vanishing Gradients more than anyone else. We’ve put out nine episodes together over the years, including panels with @jeremyphoward, @sh_reya, @eugeneyan, @BEBischof, and @charles_irl. We’re doing it again on Friday, so I went back through the archive and made this supercut. There’s Hamel deciding he hates his own skill. There’s Hamel getting frustrated with OpenClaw because the tooling had become more work than the tool. There’s Hamel, repeatedly, telling people to look at the actual data. Hamel has been saying versions of the same thing for years: stop collecting AI shit long enough to look at the thing you built. Look at your data. Read your prompts. Find the stupid failure cases. If you can’t tell why the product gave an answer, you can’t improve it. We’ll get into all of that, and more, on Friday for Stop Shipping AI Nobody Can Verify. Register to join us live, or get the recording afterwards: https://luma.com/7lng145m
RT Joe Barrow An easy but surprisingly useful trick you can learn is napkin math for model training and inference. How much work is done during speculative decoding? How fast should a DETR run at various resolutions? If you can approximate these quickly it will help diagnose slowdowns, reason about new techniques, and maybe even invent a few of your own!
RT ben hylak introducing rd-signal-2: a frontier classification model that is 1600x cheaper than GPT 5.6 Sol. free to try in @raindrop_ai, and available via a new API for training/hosting custom classifiers with Zero Data Retention.
RT Joe Barrow Good engineering is about observability. Speculative decoding accelerates LLM inference, but you're running it blind. Every rejected draft token is wasted compute, but are you looking at the drafts? This weekend I wrote specspecs to solve that!
TIL that /visualize is built in codex skill
99% of people don't know you can tell your chief of staff thread to use `/visualize` I have a pinned travel thread that tells me my travel schedule. @PhilippSpiess has done incredible work here
It’s been a long time since I’ve been excited to work through a technical book @rasbt It’s time to bring more ML back into my life
Wisprflow feels really slow, I'm motivated to uninstall it and try something new. What are ya'll using for STT? Good iOS integration is key for all the other apps
When my wife and I wfh at the same time, we both dictate constantly to AI So we can't work in the same room anymore. How is this working in the open/shared office space?