My Actors passed every test I ran. Every agent call still crashed.
I published four pay-per-event Actors to the Apify Store and exposed them to AI agents through the Apify MCP server. Every Actor passed every test I ran. Every Actor was also completely broken for the one caller I built them for.
The run failed in about four seconds. Exit code 1, zero items, no charge. And it failed before a single line of my code executed.
This is how that happens, how to catch it in CI, and which Apify counters will cheerfully tell you everything is fine while it happens. All the code is in apify-agent-path on GitHub, along with the captured output of every check.
The setup
Four Actors, all Python, all monetized with pay-per-event:
| Actor | Event | Price |
|---|---|---|
mcp-server-security-scanner | server-audited | $0.02 |
polymarket-markets-odds | market-returned | $0.005 |
sec-edgar-filings | filing-returned | $0.004 |
url-to-markdown-for-llms | page-extracted | $0.003 |
The pitch for all four is the same. An agent needs a fact, calls call-actor on mcp.apify.com, and gets structured rows back. The Actor charges per row. No human ever opens the input form.
The failure
ValidationError: meta.origin — Input should be 'DEVELOPMENT','WEB','API','SCHEDULER',
'TEST','WEBHOOK','ACTOR','STANDBY' or 'CLI' [input_value='MCP']
When an Actor is invoked through mcp.apify.com, Apify stamps the run with meta.origin = "MCP". My requirements.txt pinned apify < 3.0.0, and the 2.x SDK has no MCP member in its pydantic enum for run origin.
The call chain is what makes this expensive:
Actor.__aenter__() -> init() -> _charging_manager
Validation happens inside the charging manager, during Actor init. Three consequences:
- The run dies before the
async with Actor:body starts. None of my code runs. - It dies inside the charging manager, so nothing is billed. The caller pays $0 and gets nothing. There is no angry customer to tell me.
- Duration is about four seconds. That looks like a fast, healthy run in any dashboard that plots duration.
Why did the pin exist? It was a workaround for a dependency knot. apify 2.7.x pulls crawlee 0.6.x, which broke on pydantic >= 2.11, and browserforge 1.2.4 renamed DATA_FILES. Pinning the SDK back was the fast way out. Those bounds existed only to hold the old crawlee together, and they took the agent path down with them.
The fix is one line. The comment above it is longer than the line:
# Do NOT re-pin this back to `apify < 3.0.0`.
# That pin silently killed every AGENT-originated run: mcp.apify.com stamps
# meta.origin = "MCP", the 2.x pydantic enum has no such member, and
# Actor.__aenter__ -> init() -> _charging_manager raises ValidationError.
apify >= 3.0.0
# pydantic and browserforge are no longer constrained: those bounds existed only to hold
# the old crawlee together, and the modern SDK resolves its own.
Why every test I ran was green
Because I started every test run over the REST API.
POST /v2/acts/{id}/runs -> meta.origin = "API" -> SUCCEEDED
call-actor on mcp.apify.com -> meta.origin = "MCP" -> ValidationError, exit 1
Same Actor. Same input. Same commit. Two origins, and only one is the path a buyer takes.
This generalizes past one enum member. A green REST run proves nothing about the MCP funnel. Apify hands the run different metadata depending on how it started, and the SDK parses that metadata before your code sees it. If you sell to agents, your regression test has to come in through the door agents use.
So the fix is not "remember to test MCP." It is a script that fails loudly in CI:
# agent_buy_test.py — one cheap call-actor per Actor, over MCP.
CASES = [
("eltociear/polymarket-markets-odds", {"maxMarkets": 2, "sortBy": "volume24hr"}),
("eltociear/sec-edgar-filings",
{"companies": ["AAPL"], "forms": ["10-Q"], "maxFilings": 2}),
("eltociear/url-to-markdown-for-llms",
{"urls": ["https://en.wikipedia.org/wiki/Retrieval-augmented_generation"]}),
# Deliberately omits `mode` — see the next section.
("eltociear/mcp-server-security-scanner",
{"repos": ["https://github.com/modelcontextprotocol/servers"], "maxRepos": 1}),
]
The full script, including the MCP client it speaks through, is in the repo. Here is what it prints now that the pin is lifted:
actor status items note
polymarket-markets-odds SUCCEEDED 2
sec-edgar-filings SUCCEEDED 2
url-to-markdown-for-llms SUCCEEDED 1
mcp-server-security-scanner SUCCEEDED 1
all 4 actors run and charge through the agent path
Run it after any dependency or SDK change. A transitive bump is exactly how this class of bug comes back.
The second bug: a required field with a default
While fixing the first one I found a second, and this one is pure agent path too.
The scanner's input schema declared mode with "default": "repos" — and listed mode in the schema's required array.
Apify enforces required regardless of the default. So the sequence runs:
- An agent fetches the input schema. This is how agents learn to call your Actor.
- It reads
"default": "repos"and correctly concludes the field is optional. - It omits
mode. - Apify rejects the input before the Actor starts.
A human never hits this. The Console form pre-fills the select box from that same default. An agent hits it every single time.
A required field with a default is a contradiction, and on the agent path it costs the call. Checking for it takes ten seconds:
python - <<'EOF'
import json, pathlib
for p in pathlib.Path('.').rglob('.actor/input_schema.json'):
s = json.loads(p.read_text())
req = set(s.get('required', []))
bad = [k for k, v in s.get('properties', {}).items()
if k in req and 'default' in v]
if bad:
print(p, '->', bad)
EOF
The repo has this as schema_contradiction_check.py, which exits non-zero so it can gate a build. And the buy test above keeps mode omitted forever, so if the field ever returns to required the test goes red.
The counters that lied to me
Fixing the bug took an afternoon. Trusting the wrong instruments cost weeks. Four traps, all measured, all worth knowing before you write a monitor:
1. run.storages.datasets.default.itemCount is a snapshot, not a result. It is baked into the call-actor response. On a fast run it still reads 0 when the rows have already landed. It produced a false FAIL for me: url-to-markdown-for-llms reported items 0, while the run log said done — 1 extracted, 0 failed (charged for 1). Re-read the count from the dataset itself; treat the inline snapshot as a fallback.
2. chargedEventCounts does not exist on the run LIST endpoint. It lives only on the run detail (GET /v2/actor-runs/{id}). Read revenue off the list endpoint and every Actor you own reports a confident, wrong $0.
3. usageTotalUsd is what the run cost you in compute. It is not what the caller was charged. The two numbers point in opposite directions.
4. publicActorRunStats30Days counts your own runs too. I built an "external demand" alert on it, and it fired on my own testing. The honest field is stats.totalUsers30Days, where 1 means you. Alert on > 1, never on a subtraction between two run counters that do not mean what you assumed.
One more structural gotcha: owner runs are never billed. On my own runs, chargedEventCounts shows the Actor called charge() while accountedChargedEventCounts stays 0. You cannot self-test the paywall. You can only watch it. So build the free diagnostic accordingly: check that charge() fires and that rows land, and treat money as unobservable from your own account.
What I got wrong about the money
I would like to end with "and then the revenue arrived." It did not, and the reason is worth more than a happy ending.
For a while I told myself the crash explained $0 of revenue on a catalogue that ranked well on agent queries. Then I checked properly. totalUsers30Days and totalUsers90Days read 1 on all four Actors, and the Actors were younger than 90 days. That 1 is me. No outside caller had ever run them at all. The successful external runs I thought I saw were my own runs handed back by a counter I had misread, which is trap 4 above.
So the bug was real, the fix was real, and the causal story I built on top was wrong. The Actors were not losing sales to a crash. There were no sales to lose.
The reason sat one layer further out, and it deserves a check of its own: my Actors are in the Store index but filtered out of Store search. The Store endpoint documents a flag, includeUnrunnableActors, for Actors the platform treats as unsafe to run automatically. Mine appear only when it is set:
GET /v2/store?search=eltociear -> total 4, returned 0
GET /v2/store?search=eltociear&includeUnrunnableActors=true -> total 4, returned 4
GET /v2/store?search=eltociear&allowsAgenticUsers=true -> total 4, returned 0
Read the first line carefully. total counts them; the result set does not contain them. They are public and reachable by direct URL, and a dashboard showing a healthy published Actor tells you none of this.
I do not have a confirmed cause, and the documented ones do not fit my account. All four Actors report actorPermissionLevel: "LIMITED_PERMISSIONS", which the permissions docs name as the safe-to-run-automatically level. My identity verification is approved. It is not a cold-start threshold either: sample the default listing deep enough, at offset=3500, and it returns plenty of Actors with 0 users and single-digit lifetime runs.
The part I nearly got wrong: there are two catalogues
Having measured all of that, I drafted this section saying my Actors were invisible to agents too. Then I checked separately, and it was false.
search-actors on mcp.apify.com is the tool an agent actually calls, and it returns them:
query results our rank / actor
'github repo security scan' 7 #2 mcp-server-security-scanner
'clean markdown for llms' 10 #3 url-to-markdown-for-llms
'url to markdown' 10 #9 url-to-markdown-for-llms
'polymarket odds' 10 #9 polymarket-markets-odds
'sec edgar filings' 9 —
Ranked on four of five queries, best position #2. Meanwhile GET /v2/store returns zero of them on any query, and allowsAgenticUsers=true returns zero as well. Two discovery surfaces carry two different catalogues. The REST flag that sounds like it describes agent visibility does not match what the agent-facing server returns. Measure the surface you actually sell on.
That surface is also the tractable one. Per its own tool schema, search-actors matches on name, description, username and README. All of that is text I keep in a repo and can redeploy in minutes. Two cautions from doing it:
- The index is asynchronous and uneven. One batch of README edits moved a single Actor to #1 within minutes, and the other three did not move that day. Re-measure hours later before concluding an edit did nothing.
- The argument is
keywords. Passsearchorqueryinstead and the server does not error. It ignores the unknown key and answers with a default popularity listing, which looks exactly like "your Actors are not there."
Here is the check, because you will not stumble on either half of it by browsing:
# Half 1: are you in the index but hidden from Store search?
curl -s "https://api.apify.com/v2/store?search=YOUR_USERNAME&limit=100" \
| python -c "import json,sys;d=json.load(sys.stdin)['data'];print('returned',len([i for i in d['items'] if i['username']=='YOUR_USERNAME']),'of total',d['total'])"
# then compare against the same URL with &includeUnrunnableActors=true.
# Half 2: call search-actors on mcp.apify.com with `keywords` and look for yourself.
Both halves are scripted in the repo, as store_visibility_check.py and mcp_rank_probe.py. If the flagged number is higher than the default one, ranking is not your problem. Being in the result set at all is. And if half 1 says hidden while half 2 says #2, you have not found your answer. You have found that the question has two halves.
What I would do differently
I will not tie a bow on this. On the surface that matters, my Actors rank in the top three for the queries they were built for. Distinct users in 90 days: still 1, and that 1 is me. Being findable turned out to be necessary and nowhere near sufficient.
If I started again, I would write the agent-path test before publishing the first Actor, not after the fourth. It took an hour to build and it would have caught the crash on day one. The rest of what I would change is a habit rather than a script. Two lessons, and the second is the expensive one:
- Test the path the buyer takes, not the path that is convenient to test.
- Before explaining why a number is zero, confirm the number was ever measured. A bug
you can see is a comfortable explanation for a market you cannot.
Checklist
If you sell an Actor to agents, this is the whole post in six lines:
- Pin
apify >= 3.0.0, and never re-pin it to hold a transitive dependency together. - Add a CI test that calls each Actor through
call-actoronmcp.apify.com, not REST. - Make sure no field is both
requiredand carrying adefaultin your input schema. - Have that test omit every field an agent would reasonably omit.
- Read revenue from the run detail endpoint, never from the list.
- Alert on
totalUsers30Days > 1, never on a derived run delta.
Everything above is in github.com/eltociear/apify-agent-path, including the captured output each check produced against the four live Actors. The protocol these tools speak is the Model Context Protocol, if you want the specification behind call-actor.