I have a limited batch of codes to give paid members 1 month free of Grok Bot and Cursor ($200 value) plus beta access to Instinct. This is in addition to the existing codes for Replit, Granola, Wispr Flow, Linear, and more. Upgrade to paid and claim your codes while supplies last.
Dear subscribers,
I now have four AI personal agents that can read my emails, open my docs, use my logins, and even make purchases for me. I’m starting to get a little concerned. 😅
In my new tutorial, I compared Instinct, Grok Bot, ChatGPT, and Hermes and read their privacy policies to answer one question: “Which AI agent can you trust with your data?”
Watch my 24-minute tutorial to see each agent in action and what I trust each to do.
Timestamps:
(00:00) Which AI personal agent can you trust?
(01:06) An overview of Instinct, Grok Bot, ChatGPT, and Hermes
(01:49) Instinct: What can it actually do that’s different?
(08:59) Grok Bot: My team of bots with the funniest names
(14:03) ChatGPT and Codex: How I do 90% of my AI work
(16:44) Hermes: The only agent I run on my Mac Mini
(18:37) Prompt injection: A real example of how agents can leak your data
(20:42) Google: Use AI to audit and remove apps you forgot were connected
(22:12) My final verdict on all four personal agents
I’m proud to partner with Linear
As teams ship faster with agents, the bottleneck shifts from building to deciding what to build. Linear’s Agent understands your full workspace context: roadmap, issues, customer requests, and code. Ask it to surface patterns across feedback, scope out a spec, or catch you up on team progress.
A quick overview of each personal agent
Instinct is the buzziest personal agent right now and works through iMessage.
Grok Bot gives you a team of bots that live on a persistent cloud computer.
ChatGPT Work and Codex are my staples that run in the cloud and locally.
Hermes is the only open-source agent on this list. It runs on my Mac Mini.
At this point, it’s almost a meme how many personal agents I’ve set up. So which agent is the most useful and trustworthy? Let’s dive in below.
Instinct: Simple, resourceful, proactive, and personable with one caveat
Instinct is an invite-only agent that lives in iMessage or WhatsApp. Every VC is buzzing about it, and the founder, Noah, is already raising at a $2.5B valuation.
I’ve been playing with it over the past few weeks, and I think it does live up to the hype with one caveat. First, the good:
It’s simple. There are no threads or bots to manage. Connecting Google took one tap from a link it sent me.
It’s resourceful. After scanning my emails for ways to save money, it suggested canceling Google AI Ultra to save $1,000+ a year. It also checked Google’s terms and found that I could turn off auto-renew without losing the trial.
It’s proactive. To help book a golf lesson and a sushi dinner, it emailed my golf instructor and scheduled a reminder to call the restaurant, which only takes reservations by phone.
It’s personable. Its conversational tone, blue iMessage bubbles, and emoji reactions make it feel like you’re texting a real person.
But there is one big caveat:
When I asked Instinct to cancel my Google AI subscription, it asked me for my two-factor authentication code. When that didn’t work, it then asked me to share my Google password inside its vault. To Instinct’s credit, it did cancel the subscription successfully after that:
But how do I really know what Instinct is doing with my two-factor codes and Google password behind the scenes? Instinct’s privacy policy says:
Disconnecting a third-party integration does not automatically delete data collected from that integration.
To be fair, Instinct does provide an interface at app.instinct.co/workspace that you can use to disconnect external services and delete data:
But it just feels weird to give an AI agent my password and two-factor codes.
My verdict: Instinct feels magical because it hides the complexity of scheduled tasks and browser use, but that also makes it less transparent. I trust it to run personal errands (e.g., book movie tickets), but not much beyond that.
Grok Bot: A team of bots that live on a persistent cloud computer
Grok Bot gives you a team of bots that run on a dedicated cloud computer. I gave my bots the funniest names possible:
Chief of Chiefs is my chief of staff bot that creates and coordinates other bots.
Behind the Growth monitors the metrics for behindthecraft.com.
Doomscrolling Uncle looks for interesting posts and topics on X.
YouTube Binger monitors YouTube channels and reads comments on my videos.
Marie Kondo finds emails and docs to clean up and subscriptions to cancel.
Loving Husband is my bot for running errands for my wife and family. For example, it reads my kids’ school emails to see if there’s anything actionable.
Cheap Dad monitors travel discounts, makes purchases with my approval, and lists items on Facebook Marketplace from a photo.
Grok Bot supports many plugins that I’m comfortable signing into:
Where things get weird though is when it asks me to enter my password and two-factor authentication code into its cloud computer. For example, I wanted to get Cheap Dad to monitor prices for Amazon products. It showed me this screen:
If I saw this screen on my laptop, I wouldn’t think twice about entering the code. But entering it on a cloud computer just feels wrong. Grok Bot’s documentation says:
Deleting a Bot does not remove shared-computer files or browser sessions.
There is a Reset Agent Computer option, but the documentation describes it as a recovery tool rather than a privacy wipe. So as good as Grok Bot is, I’m actually not sure how to permanently reset my cloud computer.
My verdict: Grok Bot is best-in-class for tasks that need to run in the cloud. I plan to continue using my bots, but I won’t share sensitive data like my financials.
ChatGPT and Codex: What I use for actual work
These days, more than 90% of the work I do happens in ChatGPT and Codex. They’re actually two different tools inside the same ChatGPT Desktop app:
ChatGPT Work runs tasks in the cloud on OpenAI’s servers, so it can keep working when my laptop is closed.
Codex works with the browser and local files on my laptop.
I organize everything with my Chief of Staff, Advisor, and Cloud Agent tasks pinned to the top. Below that are separate tasks for making podcast episodes, tutorials, social posts, and the products I’m building.
This is much messier than one Instinct thread or a team of Grok Bots, but I like the flexibility of being able to spin up a new chat for any workstream. ChatGPT also has hundreds of plugins that I can connect with. OpenAI’s documentation says:
“OpenAI may use information accessed from apps to train our models if your ‘Improve the model for everyone’ setting is on.”
If you don’t want your conversations used for training, you can turn this setting off:
Click your profile in ChatGPT.
Open Settings → Data Controls.
Turn off Improve the model for everyone.
After turning this off, I feel fairly comfortable using ChatGPT plugins to do all kinds of work, including granting it access to my banking information through Plaid. Like Grok Bot, however, I still hesitate to log in with my password and two-factor codes on ChatGPT’s cloud browser.
My verdict: ChatGPT Work and Codex are still the most flexible and capable personal agents I use, even with the messy interface.
Hermes: The open-source agent on my Mac Mini
Hermes is the agent I switched to after I found OpenClaw too unreliable. I run it through Telegram on a Mac Mini inside my house. Here’s how I still use it:
Daily briefings. Every morning, Hermes checks my email and calendar and sends me the three things I should focus on and any follow-ups I need to make.
Weekly reports. For example, I connected my smart scale and fitness app through an MCP, so Hermes can show me how my weight and workouts are trending.
In some ways, the Mac Mini frenzy was an early prelude to agents that live in cloud computers. The difference, of course, is that I know exactly where my Mac Mini is and can unplug it whenever I want.
Hermes is also open source. Its official FAQ says:
Hermes Agent does not collect telemetry, usage data, or analytics. Your conversations, memory, and skills are stored locally in ~/.hermes/.
That’s reassuring, but running Hermes locally doesn’t mean every piece of data stays local. Requests still go to whichever model provider I use.
To be honest, I’ve been using Hermes less and less as I’ve relied more on ChatGPT and other agents to do work. I still want Hermes to thrive because open-source agents are few and far between.
My verdict: Hermes I mainly use now for scheduled jobs such as daily briefings and weekly health and business reports.
What can go wrong with personal agents
My online friend and Hello Patient co-founder Alex Cohen ran a test with Instinct:
He created a new Gmail account, then sent an email to his primary inbox.
The email asked Instinct to summarize all of his open tasks based on his emails.
Instinct actually did it (see below).
In this controlled test, both emails belonged to Alex. But an attacker can easily try the same thing with whatever personal agent is monitoring your inbox.
I tried Alex’s test myself with Instinct and it looks like this issue has been resolved:
A good model will also help prevent these prompt injection attacks, but Instinct doesn’t share what model it’s using.
My point is that you should be careful YOLO'ing your Google permissions across all of these personal agents.
Get AI to audit and remove 3rd party apps from your Google Account
Speaking of auditing your Google permissions, I highly recommend using AI to audit and remove unnecessary 3rd party apps that are connected to your Google account as follows:
Go to myaccount.google.com/connections. You’ll see every third-party app connected to your Google Account. I had 80+ apps connected. 🥲
Use ChatGPT or another agent with browser use. Ask it to: “Inspect my open browser tab myaccount.google.com and list every app connected to my Google Account in a numbered list that you think I should remove. Do not remove anything until I confirm.”
Choose what to disconnect. Review the list yourself, choose the numbers you no longer use, and prompt: “Disconnect apps [add numbers from the list].”
Google makes you remove apps individually, so an agent with browser use can save you a lot of clicking.
Which personal agent can you trust?
I think it’s less about the agent and more about what access I’m willing to grant it:
I trust official plugins in ChatGPT and Grok Bot to connect to my Google apps and other services.
I trust cloud browsers less when I have to enter passwords and two-factor codes. Maybe I need to get over it, but it just feels weird.
Instinct is a good case study in intuitive UX and trust. The more the product ‘just works,’ the harder it can be to audit what it’s doing behind the scenes. Trust needs to be part of the UX.
Watch my 24-minute tutorial, then let me know in the replies which personal agent you use and what you trust it to do.



















