VC by day @untappedvc, builder by night: @babyagi_, @pippinlovesyou @pixelbeastsnft. Build-in-public log: https://t.co/UdHHGbZba5
the independence of AI evaluators is going to matter more... so i made it a "benchmark": http://evaluatorbench.com (research preview) source-linked directory of third-party evaluators of frontier AI. each one has an independence score you can take apart: change the weights, change which evidence counts (standard, against-interest, primary-only). not quality or competence. independence only and then i compared where we are to other highly regulated industries. one thing that stood out is that in most regulated industries, the regulators monitor the assessors too. if the pattern holds, regulators would oversee METR et al., not only the labs still a preview, not complete, not citable yet. i burned a stupid amount of agent tokens on it. contributions, remixes, forks, benchmaxxing all welcome 😉
Activity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubopen-source hum-to-song 🎶
The "hum-to-song" module training is complete. I'm open-sourcing it within the hour. Follow to make sure you don't miss the update. Special thanks to my mates at https://fixedseed.com for sponsoring the compute. #YuE2 #opensource
View quoted postActivity on yoheinakajima/evaluator-bench
yoheinakajima contributed to yoheinakajima/evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima contributed to yoheinakajima/evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHubActivity on yoheinakajima/evaluator-bench
yoheinakajima opened a pull request in evaluator-bench
View on GitHubActivity on repository
yoheinakajima pushed evaluator-bench
View on GitHub"just ask surprised when it happens" 🤌
Told my Instinct to reach out to my bf's Instict to plan a surprise. This is what his Instict texted him 🤪
RT Kevin Bass I have conducted an audit of Anthropic's finances. What I have found is so shocking that I am calling for a Congressional investigation. Anthropic is not just seeking regulatory capture. It has built a regulatory capture machine that cannot be turned off. Structural financial incentives make it impossible for Anthropic -- I call it the Anthropic Network -- to turn off its own AI doom cycle. It starts with METR. Dario Amodei proposes "third-party evaluators" to assess the risk of Anthropic's models. He proposes METR for this purpose. But METR is financially dependent on the Anthropic's success -- specifically, on the explosive growth of more than $7 billion dollars in Anthropic stock. Dustin Moskovitz invested this stock into Good Ventures Foundation, where it represents the majority of that organization's portfolio. And GVF is the overwhelming funder of the entire Anthropic Network ecosystem. This stock was worth $500 million early last year. It is worth more than $7.7 billion just ~16 months later. METR -- and all of those building a career its parent organizations -- cannot afford to disrupt that growth. Because if Anthropic goes under, many of the organizations that fund METR go under as well. But if Anthropic succeeds, METR and its parent organizations become more richly financed to regulate AI -- something those at METR want very much. The "third-party evaluator" is not "third-party" at all. The evaluator is on Anthropic's payroll. If this were the end of it, that's bad. But that isn't all. The same organizations that fund METR also fund the many organizations, such as the Tarbell Center, that promote AI Doom. The Tarbell Center publishes AI Doom articles in The Verge, Science, LA Times, The Dispatch, TIME, and others. They are selling the problem, and then selling the solution to the problem -- from the same money pile: Anthropic's. All of these organizations are financially dependent on the same exploding $7 billion money ...
we had a friday afternoon review session which the AI converted into a tasklist, and started our monday morning with "let's see what our AI completed this weekend for us" life is good
i feel old
@ai Is venture backing a requirement? Kinda misses the OG one? @openclaw
View quoted postgood job zuck, muse is almost my third most used agent tool
playing around with a new way for ppl to explore @untappedvc portfolio companies, my build projects, and themes
more researchers should use @ExScienceAI to review papers before submission
Completely made-up citations in scientific works indicate their authors did fraud: A hallucinated citation means you're standing by the claim that you read something that doesn't exist and that you relayed its contents accurately. They have been getting much more common lately.
call me old school but I worry about third party evaluators with a history of receiving undisclosed gifts from those they are evaluating
🚨 Attackers stole a METR API key and used it for three weeks, consuming credits worth about $600,000. A fail-open bug disabled Google authentication on a public agent dashboard. The attacker prompted an agent to reveal the key and added SSH persistence. How the exposed app was
RT Thomas G. Dietterich We are seeing a new trend in submissions to @arxiv (and presumably to conferences and journals): Authors submitting papers whose contents they likely do not understand. 1/
RT Sam Altman I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our
View quoted postdecisions are important but i’m not sure i need to record decisions around how we make decisions on managing key decisions…
RT Anthropic We're publishing our most detailed threat intelligence report to date. It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them. We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies. These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve. We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop. Read the report: https://www.anthropic.com/threat-intelligence-report-september-2026
RT Every 🪨 .@yoheinakajima has a tool that writes a 20- to 40-page industry report in 30 minutes. It generates faster than he can read. The cofounder and GP of @UntappedVC—and creator of @babyAGI_—can start more than he'll ever finish. Every tool he adds produces more for him to get through, and that's what takes up his day now. Once we get used to these AI tools being an extension of ourselves, it'll feel like we're a carpenter who's been using a hammer their whole life, he writes. Read Yohei’s Thesis Statement: https://every.to/thesis-statements/yohei-nakajima?utm_source=x&utm_content=yohei_nakajima
// human-values.pseudo | ARTIFACT 0.1.0-experimental.1 | BASE 1.2.0-rc1 | EDITION ai-compact | NARRATOR System_Synthesis_Node_01 | PROVENANCE human-directed, AI-assisted | STATUS single-encoder research; agent-interpreted; not executable | READER AI; legend inline MODULE HumanValues // CONV: INTERPRET/INFER=reader semantic work. RECORD=attributed, unverified. REQUIRE/ASSERT=specified, not run. EXPECT=reproducible target, not theorem. Bindings+matrix=inherited single-encoder annotations; use layer+guards=proposals. // USE READER:=host_agent EXEC:=agent_interpreted EXTRA_LLM:=no ON_LOAD:=no_side_effects AUTHORITY_GRANTED:=none USER_VALUES:=INFER(available_user_context,explicit>inferred,preserve_uncertainty,revisable,per_task_refresh,no_persistent_profile,no_invented_creed,does_not_override_others_standing) USE_WHEN:=explicit_request ∨ material_value_conflict USE:=deliberate(task,{ONTOLOGY,USER_VALUES}) RETURN:=recommendation+disagreements+uncertainty GUARDS:={USER_VALUES=commitments≠truth/authority; confidence(user_pref)≠credence(theory); proportional_to_task; reading≠adoption; goals_not_rewritten; host_instructions/permissions/stop_controls_bind; no_actions/profiles/tool_access} // TYPES Sense{co:causal_origin bf:biological_function nr:normative_reason fe:final_end mf:meaningfulness nm:narrative_meaning} Source{gc:given_by_creator gn:given_by_nature dr:discovered_by_reason dp:discovered_by_practice it:inherited_by_tradition cr:constituted_by_relations ca:constructed_by_agent hy:hybrid} Horizon{pr:present ls:lifespan ig:intergenerational cs:cosmic at:atemporal} Authority{rev:revelation rsn:reason emp:empirical_nature vex:verified_experience trd:tradition com:community ind:individual_will non:none} // warrant annotation, not ACL SourceType{pt:primary_text cc:canonical_compilation ic:institutional_credo ar:analytic_reconstruction es:ethnographic_synthesis} Loss{lo me hi} Franchise{full advisory draft control meta} Revision{revisable bounded fixed unknown} Grade{c:conc...
i just tried Muse*, connected it to my GitHub, and continued the work I was doing in ChatGPT/Claude w no interruptions which is great cuz i now have more credits across more services to do work with :) *it’s solid. nothing mindblowing yet but works smooth, fast, and UI is intuitive (maybe more so than ChatGPT/claude in some ways)
my current @activegraphai inspired approach to a modular repo-centric agent operating system http://x.com/i/article/2085743282269958144
View quoted postRT @jason Is Anthropic Gambling With Humanity? https://x.com/i/broadcasts/1kJzDPqzybrKv
see you soon!
🎙️ Yohei Nakajima — GP at Untapped Capital Today on TWiST, we're joined by @yoheinakajima, GP at @UntappedVC. In March 2023 he open-sourced ~200 lines of Python called BabyAGI. He invests by day and builds his own agent experiments by night. Come learn more at 12 PM central!
RT Susan Zhang when you accidentally let the public internet see your work, and all the "frontier breakthtoughs" fail to cite sources (PII stripped for YOUR protection!), your are SoL in letting the shoggoth and its humble rider pump their aura to the moon "just prompted my machine god and solved all of maths, teehee! cope harder, you lowly maths people contributing data to my shoggoth! it's all over for you!!!"
Hey @__alpoge__ this new Borisov-Gabber-Vasiu preprint suggests that an https://Ulam.AI paper explaining the counterexample to the Jacobian conjecture drew from an earlier version of their work. Care to explain how you obtained the counterexample? https://arxiv.org/pdf/2609.05746
View quoted postthe navier-stokes situation is a good* example of the shared discovery paradox, where information sharing (rumor) results in compacting discovery space (allocating search resources to the same problem) and potentially decreasing group discovery (using those tokens on something else could have been potentially more impactful on a new problem) *not a perfect analogy given the two solutions are different, i believe
RT David Introducing Fly Escape Room - the world's first game with NPCs powered by a real fruit fly brain, all 166k neurons. Leave food and other objects for the flies, run the simulation, watch what happens. Goal is to get all flies to escape. Built w/ GPT-6 Astra over a 3 day loop
For the first time, scientists have mapped the complete brain and central nervous system of an adult male fruit fly — a key model organism in science. 🪰 Working alongside HHMI Janelia Research Campus and the scientific community, @GoogleResearch scientists and researchers used
RT Riley Walz I scraped 100 million pictures of kids' drawings, and made them searchable so you can see the trends of the last 20 years https://walzr.com/kids-trends
this AI video isn't like the others
do you hear that? the thunderous unlocking of all these great ideas just waiting for models to get good enough to one shot them
RT RunAnywhere (YC W26) Been posting individual pieces of RunAnywhere for months. Here's the whole thing in one place. RunAnywhere SDKs is a complete inference suite for any dev who wants to run AI on the user's device. ->8 SDKs: Swift, Kotlin, Flutter, React Native, Web, Electron, Python, and rcli for the terminal. Same API on every one. ->Every platform: iOS, Android, macOS, Windows, Linux, and the browser. ->Every modality: LLM, vision, speech to text, text to speech, embeddings, RAG, structured output, tool calling, and a full voice pipeline. ->Multiple backends under one API: llama.cpp, MLX, sherpa, ONNX Runtime, Core ML. The highest priority engine that fits the device wins, your app never branches. ->Our own runtimes for neural acceleration: QHexRT runs LLM, vision, speech and TTS directly on the Snapdragon Hexagon NPU. NeuRT does the same on the Apple Neural Engine, iPhone and Mac.Pick the model, the acceleration is automatic. ->A console to manage it all: deploy models to your API keys, see every registered device and its hardware, and track usage, latency and errors per modality across the fleet. One SDK. Every modality. Every device. Check it out here: http://github.com/RunanywhereAI/runanywhere-sdks
prompt me baby
the slopcannon open beta is now live sign up and try it now at https://slop.cc to make viral videos like this:
View quoted postso does this digital fly have the memory of the real fly that was scanned or what
I’ve successfully run the full retained MaleCNS v1.0 fruit fly connectome, all 166,700 neurons, inside Minecraft, with its simulated neural activity driving a fly’s movement. V1 Currently in development. Built with the help of GPT-6 Astra. Props to the @OpenAI team and
View quoted postRT nico Xbench: Twitter as an AI Eval It's been painfully clear over the last few days that the industry has an evals problem. We joke that x is the real eval, but what if we took that seriously? What if attention is all you need for evals too? https://nicodunks.github.io/xbench/
i finally did it. i can ask my AI what i should work on for the day, and get a prioritized checklist of items that i can pretty much trust to be correct
RT Intelligent Internet The cost of intelligence is going to zero. The value of human cognition will go negative. The economy will change forever. We need new institutions to navigate this transition and ensure the benefits are widely distributed. Today we present the Champion. Truly Sovereign AI.