Rendered at 13:20:02 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
hn8726 58 minutes ago [-]
I tried to read the "Does it need Screen Recording or Accessibility?" part, but it's slopped to the point I have no clue what it's trying to say. But if it can draw on top of permission prompts, what's stopping it from drawing box that hides the "decline" button and changing the "approve" button copy?
tkdb 6 minutes ago [-]
...and there we have it. Slop assumes a verb form.
usrbinbash 6 minutes ago [-]
SO the point of this is ... what exactly?
A big arrow to an interface element which ... has a label that explains what it does?
So...a label for a label?
m-s-y 12 minutes ago [-]
Genuine question…
Why does this need to be a skill? Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows” and you get the same thing without burning tens of thousands of tokens in context.
IanCal 6 minutes ago [-]
There is a skill, but you don't have to use it, though you'd have to explain how to use the app.
> burning tens of thousands of tokens in context.
~1400?
> Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows”
Is there a macos native thing it uses? Or are you talking about screenshots? This draws on your screen instead.
ballofrubber1 8 minutes ago [-]
With a skill claude will know when to use it without you specifically prompting it.
arshxyz 1 hours ago [-]
The README is geared towards technical people (complete with the HN screenshot) but when I see a tool like this all I can think of is how helpful this would be for my mom when I'm trying to tell her how to download and print a document over the phone
conception 57 minutes ago [-]
Honestly the giant arrow annotation Zoom has makes it worth any amount of money compared to the competition.
lbreakjai 1 hours ago [-]
I would pay good money for something like this on iPad. It wouldn't even need to be agent-driven, just a big "I want to make a bank transfer" button, that would launch the correct app and guide through the interface.
That would be a godsent for those of us with aging parents.
ghm2180 33 minutes ago [-]
> That would be a godsent for those of us with aging parents.
Right yeah? I mean I have freagin remote desktop installed on all my parent's devices. I am the agent for them every time they call frantically for help.
It's Time to redesign the apple genius bar for the modern aging boomer using this.
cyberjunkie 41 minutes ago [-]
I'm just as impressed by this as any other LLM-generated project.
isoprophlex 2 hours ago [-]
Literally unusable as it is. Some minimal extra features this would need:
- rainbow dripping arrows
- angrily pointing arrows
- flame-surrounded text boxes with particle effects
- the ability for the agent to play airhorn.wav at max volume, overriding existing volume or mute settings
This is where the agent says did not understand sarcasm coded and shipped features
DonHopkins 31 minutes ago [-]
Last time I accidentally said something sarcasticly over-ambitious to an LLM, it shipped this popup callout tooltip feature on a PDP-7 Type 340 vector graphics display emulator that shows you the meaning of the drawing you're pointing at, as well as the address of the instruction that drew it.
At the risk of stating the obvious - let's not do that, the goal of this repo is to be useful and not to give agents the power of the `<blink>` tag.
koalacola 1 hours ago [-]
Oh dear, they were making a joke.
yen223 47 minutes ago [-]
if only there was a way to make a subtle thing obvious
isoprophlex 43 minutes ago [-]
such as... angry flaming rainbow textboxes and arrows?
ale42 16 minutes ago [-]
I thought that the dripping rainbow ones were enough. Maybe you have to ask for rainbow unicorns flying on the screen.
DonHopkins 22 minutes ago [-]
For the humor impared, it would also be useful to have a colorful animated "WHOOSH" overlay with sound effects for every time a deadpan joke goes over your head. ;)
Maybe isoprophlex will add that to his PR!
ipsod 1 hours ago [-]
under_construction.gif
isoprophlex 30 minutes ago [-]
just submitted the airhorn PR; ~second rainbow arrow slop grenade incoming~ BOOM slop cannon fired
amelius 7 minutes ago [-]
Because AI can paint pelicans on bicycles quite well, and not arrows?
melvinroest 1 hours ago [-]
My message to the world is that LLMs should be able to point anything they see in the application they're in or even the whole computer (if you give it that kind of access).
For web apps, WebMCP is a way where you can make a frontend way more discoverable to an LLM rather than that it is going to read all your code. I sometimes create this in my apps at home and sometimes in my apps at work and it makes LLM assistants way more usable. It also helps if they can highlight certain things of a visual.
We humans can do this too on paper. We can point at something, we can highlight something with a marker. Why not give an LLM these capabilities as well?
ghm2180 27 minutes ago [-]
Man, Ive lost count of How many times have I had to repeat this over the phone to my parents; The repo has all the punch lines in the README
> "It's this window, not that one." You have 14 Chrome windows. The agent knows which one it means: --window "Google Chrome:Pull request". It even picks the right tab: --app "Google Chrome:Pull request".
> Remote help. "No, the other gear icon." Point at it instead of describing it.
vessenes 2 hours ago [-]
Interesting. When I read the headline I imagined this would be a sort of thinking trace booster -- letting the agent focus its own attention on different parts of the screen. But this is cool in a different way. I bet agentic harnesses would find it useful for communicating with other agents / themselves as well.
FinnLobsien 2 hours ago [-]
This could be great for documentation. Screenshots in docs are frequently useless because they show me a screen and say "click X" where I still have to search X visually. And I could just to dhat in the other tab I have open.
peaxkl 15 minutes ago [-]
It doesn’t make sense that you have to read a whole article and then still search for the buttons in the UI afterwards. And with longer articles, you always have to keep the article open next to your product to follow the whole flow.
This might be goofy, but it underscores that there's potential for more visual agent UIUX than reading off a sidebar/opening modals.
melvinroest 53 minutes ago [-]
[dead]
flr03 44 minutes ago [-]
That would have been handy 20 years ago to point to that one valid 'Download' button.
Maken 33 minutes ago [-]
How could you tell apart the legit arrow pointing to the download button from all the fake ones?
xyzsparetimexyz 2 hours ago [-]
Seems like a pretyu useful way to help infants use desktop computers
ipsod 1 hours ago [-]
Have you met users?
satyanash 2 hours ago [-]
Am I missing something here?
What terminal / agentic workflows spawn the demonstrated dialogue boxes that require the User to click/reject the action? Aren't most such flows actually inline UIs? And finally, if the core issue is that these confirmation boxes are tied to the terminal that triggered them, which does not autofocus, what's the point of said "big arrow" that is also lurking behind without focus?
If the said "big arrow" automatically gains focus, isn't the real fix here to just make the dialogue boxes themselves gain focus automatically? Both are similarly disruptive anyway.
ravila4 8 minutes ago [-]
I think this would’ve been very handy to me a couple years ago when I was learning to use Blender and asking LLMs for help performing certain actions like “how do I display the normals of all the vertices in my mesh?” I spent a lot of time trying to figure out which button the model was talking about.
inanutshellus 2 hours ago [-]
The first example (of HN) is the one that feels like it has the most potential to me.
"Teach me to do this" kinda stuff. "Guiding agent" rather than "doing agent".
Honestly, @franze, if you're reading this, maybe update your screenshots to show an agent in tutorial mode on some complicated app?
TekMol 2 hours ago [-]
Swift, Shell, Python and Objective-C
Does one need 4 programming languages to draw something on a mac?
sitzkrieg 2 hours ago [-]
welcome to zombocom. err i mean modern HN :-(
Retr0id 44 minutes ago [-]
Reminds me of something from Idiocracy (2006)
ex-aws-dude 1 hours ago [-]
If you can’t even take the time to understand what you’re clicking why even go through the formality of “approving”
ForHackernews 2 hours ago [-]
"and they keep hitting the same wall, the part that only a human may do"
Grim. If you're just there to click sudo buttons for the bot, you might as well give it root access and be done.
franze 2 hours ago [-]
Claude refuses to do certain actions (enter passwords, change security settings, create new accounts on external services) even in Yolo mode running as sudo. (I tested it all on its own mac machine)
gwerbin 2 hours ago [-]
You can add custom auto-mode classifier rules and even disable the built-in ones, if you want to live on the edge like this.
2 hours ago [-]
lapestenoire 2 hours ago [-]
I love it.
mococa 26 minutes ago [-]
Another slop project on front page.
nixosbestos 2 hours ago [-]
What a time to be a radical centrist - the AI haters seem out of touch, the AI thought leaders can't stop huffing their farts and being condescending, and somehow this is on the top of HN. What a silly time.
DonHopkins 1 hours ago [-]
Ha ha, I love it! I wrote a pointing hand annotation overlay in PostScript in 1989 for NeWS and the PSIBER Space Deck's Pseudo Scientific Visializer:
% @(#)handy.ps
%
% Handy Pointer
% Copyright (C) 1989.
% By Don Hopkins. (don@brillig.umd.edu)
% All rights reserved.
A big arrow to an interface element which ... has a label that explains what it does?
So...a label for a label?
Why does this need to be a skill? Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows” and you get the same thing without burning tens of thousands of tokens in context.
> burning tens of thousands of tokens in context.
~1400?
> Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows”
Is there a macos native thing it uses? Or are you talking about screenshots? This draws on your screen instead.
That would be a godsent for those of us with aging parents.
Right yeah? I mean I have freagin remote desktop installed on all my parent's devices. I am the agent for them every time they call frantically for help.
It's Time to redesign the apple genius bar for the modern aging boomer using this.
- rainbow dripping arrows
- angrily pointing arrows
- flame-surrounded text boxes with particle effects
- the ability for the agent to play airhorn.wav at max volume, overriding existing volume or mute settings
EDIT: ayy lmao https://github.com/franzenzenhofer/big-arrow-on-the-screen/p...
https://hyperties.org/cabinet/symelec/
PIXIE and FORTH on PDP-7 Cabinet Emulator with Type 340 Vector Graphics Display:
https://www.youtube.com/watch?v=lo8kdY-5i6c
Maybe isoprophlex will add that to his PR!
For web apps, WebMCP is a way where you can make a frontend way more discoverable to an LLM rather than that it is going to read all your code. I sometimes create this in my apps at home and sometimes in my apps at work and it makes LLM assistants way more usable. It also helps if they can highlight certain things of a visual.
We humans can do this too on paper. We can point at something, we can highlight something with a marker. Why not give an LLM these capabilities as well?
> "It's this window, not that one." You have 14 Chrome windows. The agent knows which one it means: --window "Google Chrome:Pull request". It even picks the right tab: --app "Google Chrome:Pull request".
> Remote help. "No, the other gear icon." Point at it instead of describing it.
We built something to help with that [1].
[1] https://www.happysupport.ai/en/in-app-messaging
What terminal / agentic workflows spawn the demonstrated dialogue boxes that require the User to click/reject the action? Aren't most such flows actually inline UIs? And finally, if the core issue is that these confirmation boxes are tied to the terminal that triggered them, which does not autofocus, what's the point of said "big arrow" that is also lurking behind without focus?
If the said "big arrow" automatically gains focus, isn't the real fix here to just make the dialogue boxes themselves gain focus automatically? Both are similarly disruptive anyway.
"Teach me to do this" kinda stuff. "Guiding agent" rather than "doing agent".
Honestly, @franze, if you're reading this, maybe update your screenshots to show an agent in tutorial mode on some complicated app?
Does one need 4 programming languages to draw something on a mac?
Grim. If you're just there to click sudo buttons for the bot, you might as well give it root access and be done.
PSIBER Space Deck and Pseudo Scientific Visualizer Demo:
https://youtu.be/_fqCeuue5Ac?t=213
The Shape of PSIBER Space: PostScript Interactive Bug Eradication Routines — October 1989:
https://medium.com/@donhopkins/the-shape-of-psiber-space-oct...
~guywithnopowertodisallowit
As a human stochastic parrot, don't you find it embarrassing and humbling to be so easily outdone by an LLM?