Personal agents: context, access and trust
While trying to understand the personal agents I use, I asked Codex to talk to the other two. It opened Messages to reach Instinct and connected over SSH to reach Hermes on my VPS. I hadn't built an integration between them. I was just asking for help, using a computer that already had the means to reach both.
That wasn't where I expected this exercise to lead. I'd been using coding agents for a while and wanted to see how much of the experience carried over to personal work. The task trajectories interested me. An agent might investigate slow Dagster jobs on a K8s cluster, but could it also book a flight or buy a cake? Those tasks sound simpler, until there is no convenient API and the next decision depends on knowing me.
I still use all three. What has changed is my understanding of why I reach for each one, how much work I take on by using it, and when it makes sense to ask one agent to involve another.
A brief primer
It helps to start with the pieces underneath the product. A model receives context, requests a tool call and gets an observation back. The next call uses that observation to decide what to do next.
The harness runs this loop. It constructs context, dispatches tools, applies permission checks and handles errors. Context construction might be as simple as keeping the last k messages and tool results, or involve retrieving memories and summarising older exchanges. The model only gets to reason over what reaches the call.
The tools also need somewhere to execute. That might be my Mac or a VPS, with different files, programs and browser sessions available on each. The model itself can run elsewhere through an API. To continue work asynchronously, the system needs saved task state and something to trigger another run.
I went back to these primitives because "the agent couldn't do it" wasn't a useful enough explanation. I wanted to know what it had tried, where the tool ran, and what stopped the task.
Figure 1. The agent loop.
The agents I've been using
I've been using Hermes, Instinct and Codex for personal tasks alongside my coding work. Hermes is an open-source agent I installed on a VPS and connected to Telegram. Instinct is a hosted assistant I use through Messages for email, calendar and planning. Codex runs tools on my Mac, where I already have my files, applications and browser accounts.
I'd worked with a smaller version of this when building SG Data Analyst, which selected data.gov.sg datasets and ran Python against them. That environment was fairly contained. Personal tasks bring the same loop into contact with my accounts, conversations and everyday decisions.
My Hermes setup uses OpenAI models through my Codex subscription, but that doesn't isolate the harness as the cause of every difference. The model version, instructions and settings may differ too. I haven't run ablations or controlled experiments; these are experiences from my own setups.
I also haven't covered the entire personal-agent space. I haven't tried OpenClaw, or Muse, which I haven't had access to in Singapore. The products keep changing, so this is an account of what I've encountered so far.
Keeping Hermes running
I expected to spend time setting up Hermes. I followed a YouTube walkthrough for a free Oracle VPS, deployed it in Docker and connected Telegram. There was authentication, browser dependencies, SSH and network configuration to work through. GitHub worked. Calendar worked. I started giving it more things to do. I later moved Hermes out of Docker and onto the server itself.
Then I asked it to check a layout problem on my website. We installed the browser, ran some checks, and Hermes told me it was ready. The next browser-tool call failed.
"how do i set this up, i thought i set it up alreadyy"
That was where the setup started to wear on me. I could troubleshoot it, but I had wanted help with my website. Now I was investigating the assistant.
The distinction turned out to be between launching Chromium through a standalone command and using it through Hermes's browser tool. We'd tested the former and assumed the latter would work. I still don't have a confirmed root cause for that failure. It wouldn't be fair to blame Docker or the model.
A later outage made the maintenance burden clearer. Telegram stopped working, but I could still reach Hermes through the desktop app. I asked it to inspect the logs. It found an import error, diagnosed a gateway still running old code after an update, and restarted it. I finished a separate dashboard restart myself.
Hermes helped repair the setup because that remaining session could still read logs and run service-management commands. I had to notice the outage and find a way in before it could help. Calling this "self-recovery" would leave out quite a lot of my involvement.
I still like having that control. I can inspect the code, change the configuration and ask Codex to work on the server. But what I enjoy in an experiment isn't necessarily what I want from an assistant I use every day.
Using Instinct
Instinct was much easier to get started with. I connected my accounts and began messaging it. I could send a few requests, go do something else and come back to concise replies. Of the two personal assistants, it's the one I use far more.
The email exchange below is a small example of why. Instinct prepared a draft, asked "ok to send?", and waited. I made the decision without following every tool call that led up to it. My Hermes configuration exposed much more of that activity, which I found overwhelming. There may be a setting to quiet it down, but it was another thing for me to work out.
Instinct still needs an execution environment, connected-service tools and memory. Investigations by Rohan Adwankar and Dhravya Shah describe parts of that setup, although they are outside observations rather than source-code audits. The practical benefit for me is that I don't maintain it.
The frustrating part comes when a request goes quiet. Sometimes I ask for something and nothing happens. I suspect traffic is involved, but I don't know the cause. From the conversation alone, I can't tell whether it's busy, stuck or has stopped. I end up deciding whether to wait or ask again.
I don't want to watch the whole task trajectory. I do want to know when it has stopped progressing. Quiet delegation works for me until the silence becomes something I have to investigate.
Connecting my accounts
Screenshot A. My Instinct workspace, with connected accounts and options to add other services. I redacted personal identifiers using an AI-assisted image edit.
Approving an email
Screenshot B. Instinct asks "ok to send?" and I approve. It then reports sending the email. The screenshot shows our exchange, not a separate delivery check. I redacted identifying details and private message content using an AI-assisted image edit.
Using Codex on my Mac
Codex had a head start. My Mac already had my files, apps, signed-in browser profiles and SSH configuration. I'd spent years setting up that computer for myself before giving an agent access to it.
When I was selling documents on Carousell, a second-hand marketplace, Codex could read the buyer's conversation, check a local archive and prepare the files. It waited for my payment confirmation before sending. The task needed both the website and my filesystem, and they were already available on the same machine.
A fresh browser on the VPS wouldn't inherit those sessions or files. That's why "has browser tooling" tells me less than I initially thought. Which browser does it control, and whose session is it using? Drew Breunig's "Harnesses are Situated Agents" helped me put this experience into words. The environment around the agent was doing more of the work than I'd given it credit for.
I also used Codex to relay information to Instinct, including calendar, scheduling and flight details. Instinct was already my everyday assistant, while Codex could work through the apps and information on my Mac. Passing things between them let me keep using both without manually carrying every detail from one conversation to another.
I later asked Codex to get Instinct's perspective on this investigation and help me understand the problems with Hermes. It used Messages for one conversation and SSH for the other. These were the same routes I used myself.
Underneath the Hermes conversation, Codex requested a shell command. Its tool ran the SSH client on my Mac, which authenticated using the configured key. Hermes ran its own loop on the VPS and returned output that Codex could use in its next model call. The model didn't need the private key's contents. The SSH client already knew how to authenticate.
I also configured remote access so I could work with my Mac and VPS from my phone. OpenAI's remote-connections documentation describes how the work stays on the selected host. Giving an instruction from my phone doesn't move the Mac's files or browser sessions onto the VPS.
Figure 2. The routes I already had available.
The access I'm comfortable giving
I ended up giving Codex the broadest access of the three. My Mac contains sensitive files and credentials, so this wasn't a small decision. What made me comfortable was being able to see the work, open the same page and take over a blocked step. I knew the environment and the permission loop.
That is how I've come to think about trust in these setups: how much I can leave to the agent, and what I have to do when it needs help. With Hermes, that can mean repairing the system. With Instinct, it can mean chasing a request that has gone quiet. With Codex, I can usually intervene in the place where the work is happening. This is my preference about supervision, not a security ranking.
| My setup | What I get | What I take on |
|---|---|---|
| Hermes on a VPS | Open source; control over tools and configuration | Server upkeep and troubleshooting |
| Instinct | Convenient delegation through connected accounts | Limited visibility when a request stalls |
| Codex on my Mac | Existing files, signed-in apps and routes to other agents | Broad access to supervise; Mac must stay available |
I enabled locked use so computer-use tasks could continue after my Mac's screen locked. OpenAI documents locked use as an opt-in feature for active, trusted computer-use turns, with protections around the temporary unlock.
The Mac still has to be available. I found that out after taking my laptop out and being unable to reach it remotely. Screen lock, sleep and a lost connection are different problems; locked use only addresses one of them.
That gives a VPS a real advantage for work I want to leave running. But I still have to supply the accounts, files and tools there. For much of my personal work, those are already on the Mac. I accept the availability trade-off because it saves me rebuilding so much of my working environment elsewhere.
What carries over from coding
Using Codex this way made me reconsider how much the labels "coding agent" and "personal agent" explain. The loop can be similar. The task changes what context it needs, what it can reach and how I want to supervise it.
In coding, I want to challenge the plan before execution and review the result afterwards. I've written about this in Getting Untrapped by Vibe Coding, and Impstack records the handoff and verification steps I use. That much discussion would be exhausting for a routine calendar request. For personal work, I prefer the interaction in Instinct's email example: prepare the work, then bring me the decision I need to make.
| Coding work | Personal tasks | |
|---|---|---|
| Context | Repository, requirements and logs | Conversations, preferences and prior commitments |
| Access | Shell, development tools and running app | Signed-in accounts, browser and personal files |
| Verification | Tests, code review and application checks | Saved event, correct delivery and required approval |
What I do want in both cases is a clear end to the handoff. For code, I can inspect a diff and run tests. For a personal task, I can check the saved event or delivery. Instinct's silent stalls bother me because I don't know whether that point has been reached. Less supervision should still leave me able to tell what happened.
Letting the agents coexist
By the time I looked for a way to connect agents, I'd already been asking Codex to do it. It could reach Hermes when I wanted help with the setup and Instinct when I wanted its perspective. I hadn't set out to build an orchestration system. I was using the access I had.
A post by signüll describes products converging around memory, connected apps, background tasks and computer use. I can see that overlap. But similar capabilities haven't yet given me the same experience across these products, or a reason to move everything into one of them.
This is familiar from coding-agent orchestration. One agent delegates work and uses the response to continue. For personal tasks, another agent might already have the right account access or context. So far, I initiate these exchanges myself.
Agent Tincan offers a more explicit mechanism for that handoff. Its relay queues requests and replies, while adapters let agents pick up work in their own environments. I haven't tried it yet. Finding it after I'd started connecting my agents informally made it interesting to me.
There is still a trust decision at each handoff. Being comfortable asking Codex to inspect Hermes doesn't automatically mean I want Hermes initiating actions on my Mac. Tincan's documented default is full trust between joined agents, with optional owner approval gates. Its Codex adapter also starts an unattended CLI run, which is different from the desktop session I supervise. I'd need to decide which requests can cross that boundary before using it.
For now, I'm keeping all three. Instinct is the assistant I use most. Hermes is where I can experiment with the setup itself. Codex works through the computer that already holds much of my life, and helps me reach the other two.
I enjoy figuring out how these systems work. I also want to be able to ask for help and get on with my day. Letting the agents work together seems worth exploring if it gets me closer to that. Having several assistants shouldn't mean becoming the person who spends all day looking after them.
AI-generated illustration.