Agentic Index
AI Library vs Pickaxe (2026)
Both are no code platforms for building and running agents and they aim at opposite buyers, at 6 and 10.5 of 14. That verdict is the Agentic Index coverage score, graded from each vendor's own published materials.
AI Library is enterprise shaped: build and run agents across the full delivery lifecycle, grounded with retrieval and human oversight, free to start with enterprise delivery priced by sales. Pickaxe is creator shaped: a visual builder with a vector knowledge base, 500 plus integrations, 50 plus models, one line embeds, branded portals, Stripe billing and access control, free Starter then 29 dollars a month. Human oversight and delivery lifecycle against embeds and billing.
This comparison is published by Agentic Index, an independent agentic AI vendor research platform. AI Library and Pickaxe are each graded against the same 14 capability Agentic Index taxonomy, from the vendor's own public materials under the Agentic Index verification standard, alongside 969 researched vendors. No vendor pays for placement and no vendor has reviewed this page. How this evidence is graded
Choose AI Library if
- Human oversight built into the lifecycle is what your organization requires.
- Enterprise delivery with a vendor accountable for it is how you buy software.
- Grounding with retrieval across the delivery lifecycle is the scope you need.
Choose Pickaxe if
- Documented coverage is broader and you want to ship something to users quickly.
- Stripe billing and branded portals mean the agent can be a product, not just a tool.
- Twenty nine dollars a month with a free Starter tier is the whole commercial decision.
| At a glance | AI Library | Pickaxe |
|---|---|---|
| Category | Agent builder | Agent builder |
| Entry price | Free to start; enterprise delivery priced by sales | Free Starter plan; Gold at twenty nine dollars per month and Pro at ninety seven dollars per month, plus credit based usage billed at cost. |
| Free / trial | — | Free Starter plan |
| Pricing confidence | public partial | public exact |
| Feature | A AI Library |
P Pickaxe |
|---|---|---|
| Action & orchestration | ||
|
Integrations & Tool Calling Ability to connect agents to real systems through native integrations, OAuth-authenticated actions, custom tools, APIs, webhooks, or MCP-compatible tools. |
Partial
Stands at P, re-based off the July basis that admitted a named connector list was not retrieved. It still is not, but the shape of what exists is now documented. WHAT IS DOCUMENTED IS CAPABILITY BREADTH, NOT INTEGRATION BREADTH, and that distinction is what holds the grade. UTILITIES give agents special skills: web search, news search, web page scraping, and parsing of PDFs and handwritten text. Knowledge bases read from uploaded documents, files on the internet, and THE CUSTOMER'S DATABASE. Forms collect structured data and trigger follow-up work. Those are real tool-calling capabilities and a genuine route. WHAT THIS AXIS MEASURES IS BREADTH ACROSS CLASSES, and no named connector to any application class appears anywhere across three passes. There is no CRM, ERP, ticketing, messaging or calendar integration named, no connector catalogue, and no count. The vendor references ENTERPRISE INTEGRATIONS as a category its coding agent assembles from, without naming one. AI Library MCP, announced around May 2026, is described as eliminating fragmented integrations by giving agents structured access to tools, data and workflows across search, parsing, storage, retrieval and APIs. Read carefully, that list is the utility set rather than a set of application connectors, so it strengthens the same half already credited rather than adding the missing one. Its outbound direction is credited on Ext. THE HONEST READING for a buyer: this platform reaches the open web, documents and databases well, and whether it reaches their specific business applications is a question they would have to ask. That is precisely Partial. |
Full / Explicit
Stands at F, now anchored on the vendor's own browsable Actions directory rather than on a summary sentence. BREADTH ACROSS CLASSES IS DOCUMENTED AND VERIFIABLE. The Actions directory at pickaxe.co/actions is an enumerated catalogue rather than a claimed count, spanning CRM with HubSpot, Salesforce and Attio, communication with Gmail and Slack, productivity with Notion, Google Calendar and Google Sheets, developer tooling with GitHub, and generation with DALL-E, alongside Composio for onward tool authentication. Those are genuinely different integration classes, which is what separates this from a deep catalogue inside one ecosystem. TOOL CALLING IS REAL AND BIDIRECTIONAL. The vendor's framing is that Actions let YOU AND YOUR USERS TAKE ACTIONS AND TRIGGER WORKFLOWS WITH THE TOOLS IN REAL TIME, so the agent both reads and writes, and the end user of a deployed agent inherits that reach rather than only the creator. THE MCP DIRECTORY IS THE PART THAT SCALES IT. Pickaxe publishes MCP servers as Actions, so any MCP-exposed service becomes agent-callable without Pickaxe building a connector for it. That is the inbound direction, an agent calling out through MCP, and it is credited here rather than on Ext, where the outbound direction sits. Custom Actions plus Zapier and Make webhooks close the ceiling for anything not in the catalogue. ONE COMMERCIAL LIMIT WORTH RECORDING because it affects what a buyer actually gets: the Gold tier caps Actions at three per agent, with the cap removed on Pro and Business. A capability gated by plan is still the capability, but the entry tier is materially narrower than the 500-plus figure suggests. |
|
Workflow Orchestration Ability to sequence, branch, retry, route, and combine deterministic workflow nodes with autonomous agent steps. |
Full / Explicit
Stands at F, and the basis moves off the delivery-lifecycle prose it rested on, which described how the company works rather than what the platform does. WHAT CARRIES IT NOW IS FIRST-PARTY AND PRODUCT-SIDE. The vendor's own SDK repository describes itself as a framework to build AI agents AND MULTI-AGENT ORCHESTRATIONS, which is the axis stated plainly by the vendor about its own code. The company page corroborates with a named solution shape, a MULTI-AGENT ORCHESTRATION LAYER WITH APPROVALS AND ESCALATION LOGIC. MULTI-STEP EXECUTION IS DOCUMENTED THROUGH FORMS, where structured responses trigger follow-up tasks. That is a documented handoff from one step to the next based on collected data, which is a sequencing mechanism rather than a claim about one. AI Library Code is the second half: a coding agent that generates and assembles solutions from reusable platform components and enterprise integrations, so the platform composes the system and deployed agents run it. The vendor's framing, AI LIBRARY CODE BUILDS THE SYSTEM, AI AGENTS RUN THE SYSTEM, is a clean statement of that split. Scale corroborates rather than carries: the vendor reports over thirteen million agent actions in production as of May 2026, and describes autonomously reconciling invoices and resolving support tickets end to end. WHAT IS NOT DOCUMENTED, and the reason confidence is medium rather than high: no control-flow vocabulary appears anywhere reached. No branching, looping, conditional or parallel construct is named, and no orchestration configuration surface is described in the API documentation. The grade rests on documented multi-agent composition and step sequencing, which is present, rather than on graph expressiveness. |
Full / Explicit
Stands at F, with the multi-agent mechanism named more precisely than the July basis managed. SUB-AGENT ROUTING IS THE LOAD-BEARING FACT: an agent routes to specialised sub-agents in a WATERFALL configuration for multi-step tasks. A waterfall is an ordered fallback chain rather than a parallel fan-out, which is a modest but real orchestration primitive and an honest description of what a no-code builder for solo creators would actually ship. MULTI-STEP EXECUTION IS DOCUMENTED THROUGH THE WORK. Post-submission workflows fire after a conversational form completes, Actions chain into connected tools, and the action set spans PDF generation, structured data collection through conversational forms, and code execution. Sequencing across generation, collection and external writes is orchestration whether or not a canvas exists. THE ABSENT PRIMITIVE IS THE HONEST LIMIT AND IT IS THE SAME ONE AS ELSEWHERE IN THIS LANE. No branching, looping, conditional or parallel construct is named on any page reached, and there is no visual flow canvas. Control flow is the model's decision plus a waterfall, which suits the product's audience and would not suit a process that must run the same way every time. CONFIDENCE IS MEDIUM RATHER THAN HIGH for that reason: the grade rests on documented multi-agent routing and step sequencing, which are present, rather than on an inspectable orchestration surface, which is not. This is the cell most likely to read differently to a buyer coming from n8n or FlowX than the letter suggests, and worth a note on any comparison page pairing them. Code execution is recorded here as an action type and is deliberately refused on Comp, per the standing convention that executing code is not operating software. |
|
Triggers & Channel Coverage How agents wake up and where they work: schedules, webhooks, message events, CRM events, inbox events, chat, email, voice, and collaboration tools. |
Partial
Stands at P, but the July basis said triggers and channels WERE NOT ENUMERATED and they now are, so the grade is held on evidence rather than on absence. THE CHANNEL CLASS IS DOCUMENTED AND IS TWO WIDE. The API documentation states a conversation with an agent can be triggered AS A TEXT CHAT OR A VOICE CHAT. Voice as a first-class conversation mode on a pre-seed platform is more than most peers at this stage document, and the vendor publishes starter chat application repositories in Next.js and Angular, so the embedded surface is real rather than notional. THE INBOUND API CLASS IS ALSO MET, through the REST API and Python SDK now credited on Ext. An agent is invocable programmatically from anything the customer runs. A THIRD ROUTE IS DOCUMENTED AND IS THE MOST INTERESTING ONE. Forms let agents collect structured information, and the documentation states you can TRIGGER FOLLOW UP TASKS BASED ON RESPONSES YOU GET. That is work reaching an agent because a form was completed, which is genuinely event-driven and specific to how this product is used. WHAT HOLDS IT AT PARTIAL is the schedule class, which is absent. No timer, cron, recurrence or scheduled execution appears anywhere across three passes, and no inbound webhook or external event subscription is documented beyond the API itself. For a platform selling invoice reconciliation and reporting, recurring execution is the obvious missing route, and its absence from the documentation is more likely a gap in what is written down than in what exists. Confidence rises to medium because the classes present are now named first-party rather than inferred. |
Full / Explicit
Stands at F and is the best-evidenced channel cell reviewed this session, because coverage is stated per channel with the delivery mechanics named rather than as a list. SIX ROUTES ARE DOCUMENTED INDIVIDUALLY: website embeds in inline, floating, popup and iframe forms; branded portals on custom domains with user accounts and access groups; a Completions API; email; Slack; and WhatsApp, with third-party review adding Telegram and Discord. THE EMAIL DETAIL IS THE ONE WORTH CARRYING. The vendor documents running AN AGENT ON A REAL EMAIL ADDRESS WITH THREADED MEMORY AND ATTACHMENTS. Threading and attachments are the two things that make email genuinely usable as an agent channel rather than a notification pipe, and most platforms claiming email coverage do neither. THE EVENT CLASS IS COVERED SEPARATELY IN BOTH DIRECTIONS. Actions fire workflows in connected tools in real time, and outbound webhooks notify Zapier, Make and other endpoints. The Completions API gives an inbound programmatic route, and authenticated embeds handle session handoff from a host application. WHAT IS ABSENT IS THE SCHEDULE CLASS. No timer, cron or recurrence appears on any page reached. For a product whose agents are summoned by an end user rather than run unattended that is a coherent omission, and channel plus event coverage carries the grade comfortably without it. A property specific to this vendor: because portals carry accounts and access groups, a channel here also carries identity, so the same agent recognises a returning user across visits. That is what makes the persistent per-user memory credited on Mem usable rather than theoretical. |
| Knowledge & context | ||
|
Knowledge Grounding & RAG Ability to ground agent behavior in company data through document ingestion, retrieval, external knowledge APIs, semantic search, or RAG layers. |
Full / Explicit
Stands at F, moved off LinkedIn and onto the API documentation, and it is now among the better-evidenced cells on this record. THE KNOWLEDGE BASE IS A FIRST-CLASS OBJECT, which is what the persistence line asks for. The documentation is explicit that an agent by default draws only on the model's own knowledge, and that grounding on customer data requires CREATING A KNOWLEDGE BASE. Making it a separate addressable object rather than a context-stuffing step is the distinction between a maintained retrieval structure and per-request assembly. THE SOURCE BREADTH IS THE STRONGEST PART. A knowledge base is fed from documents the customer uploads, FILES ON THE INTERNET, and THE CUSTOMER'S DATABASE. That third route matters disproportionately: most small vendors on this axis ingest uploads only, and reading a live database means grounding on the current state of a system of record rather than on a snapshot someone remembered to re-upload. The SDK corroborates the binding mechanically: an agent carries a knowledge_id, and uploads target that identifier, with txt, pdf, pptx, docx and xlsx accepted. So the index persists and belongs to the agent rather than to a conversation. Utilities extend grounding at retrieval time with web search, news search, page scraping, and parsing of PDFs and handwritten text. Handwriting parsing is an unusual inclusion and points at the document-heavy operational work the vendor sells into. One limit recorded: no chunking, embedding, re-indexing or retrieval configuration is described on the documentation index, and refresh behaviour for database-backed knowledge bases is not stated. |
Full / Explicit
Stands at F and clears the persistence line comfortably, which is the discriminator this axis turns on. AN INTEGRATED VECTOR DATABASE IS THE MAINTAINED STRUCTURE. The knowledge base is a first-class object the creator builds once and the agent queries thereafter, not context assembled per request, and it is shared at workspace level so several agents ground on the same corpus. SOURCE BREADTH IS THE STRONGEST PART: more than 14 content types including PDFs, websites, Notion, Google Drive and video. Video ingestion is uncommon in this lane and matters for the buyer Pickaxe serves, since coaches and consultants hold much of their expertise in recorded material rather than documents. THE FRESHNESS MECHANISM IS WHAT MOST PEERS LACK. AUTOMATIC DAILY SYNCS mean a knowledge base connected to Notion or Drive tracks the source rather than freezing at upload. The usual failure of a document-index product is a corpus that silently goes stale, and a documented sync cadence is the direct answer to it. PER-FILE CONTEXT INSTRUCTIONS ARE THE DETAIL WORTH CARRYING. A creator annotates individual documents with guidance on how the agent should use them, which is retrieval steering at the document level rather than one global prompt. For a knowledge base assembled from mixed material, a price list, a policy, a transcript, telling the agent what each is for is a meaningful quality lever. The CLI syncs knowledge bases programmatically, so the corpus is maintainable from code as well as the web app. What is not documented is chunking strategy, embedding model, retrieval configuration or citation behaviour in responses. |
|
Memory & State Persistence Ability to persist context across a run, conversation, workflow, user, team, or longer-term memory layer. |
Partial
RESOLVED FROM UNRESOLVED, grade held at P but now founded rather than guessed. The July basis asserted that agents RETAIN MEMORY TO EXECUTE MULTI STEP TASKS while admitting no store was detailed, which is a claim with nothing behind it. The published Python SDK gives a real basis. THE ROUTE THAT WORKED, after six failed retrievals on the docs tree in the previous pass, was the SDK rather than the documentation. A client library's quick start enumerates the resources a platform exposes, which is a different and often better surface than prose. THE CHAT INTERFACE IS STATELESS AND THAT IS THE FINDING. The documented call is agent.chat(messages=[...]) taking a full array of role and content pairs supplied by the caller. There is no thread identifier, no session object and no conversation resource. The client holds the history and passes it in on every call, which is the completions shape rather than the assistant shape. Under the standing question of whose mechanism it is, the conversation buffer here belongs to the customer's application, not to the platform. THE DOCUMENTED RESOURCE SET IS agent.create, files.upload, knowledge_base.get_status and agent.chat. No memory, thread, session or state resource appears in it. WHAT PERSISTS IS KNOWLEDGE, NOT MEMORY, and the two are kept apart deliberately. Agents carry a knowledge_id and files uploaded against it persist in a knowledge base with a queryable status. That is a maintained retrieval structure and it is graded on Know. Counting it here would be the double-count this record otherwise avoids. WHY THIS IS NOT DOWNGRADED TO N despite the evidence pointing that way. What I read is a quick start, not a complete API reference. The src/ailibrary module tree, which would enumerate every resource, is robots-disallowed on GitHub, and docs.ailibrary.ai has failed retrieval across two sessions. Asserting that a capability is absent from a partial enumeration is the error this lane has corrected repeatedly, so Partial stands with the tension named: N is defensible on what is visible, and a full API reference would settle it in one read. |
Full / Explicit
P>F, and it clears both limbs of the 31 August ruling rather than only the state one, which is unusual. THE STATE LIMB IS STATED IN THE VENDOR'S OWN WORDS. The features page says the platform gives agents PERSISTENT MEMORY ACROSS SESSIONS. That is the ruled Full condition verbatim, and the July basis missed it by grading from the Configure tab's context-window token allocation, which is a within-session concern and correctly reads as Partial on its own. THE ACCUMULATION LIMB IS MET TOO AND IS THE BETTER FINDING. A USER MEMORIES tab lets the operator CONFIGURE TRAITS THAT AGENTS WILL COLLECT AND REMEMBER ABOUT USERS ON A USER-BY-USER BASIS, naming a user's name, role and preferences as examples. Memory that the agent gathers over time, keyed per end user and shaped by operator-declared traits, is the accumulation route, and configurable extraction targets are a more deliberate design than a generic memory store. THIS IS ALSO ADDRESSABLE, WHICH THE RULING REQUIRES. Memory is exposed as an API resource rather than an internal effect: a community MCP server built against the public API calls memory_list and memory_get_user to read per-user memory. That is third-party code and carries nothing on its own, but it demonstrates the memory objects are real, named and reachable through the documented API. WHY THIS MATTERS FOR THE PRODUCT rather than being a checkbox: Pickaxe agents are sold to end users through branded portals with accounts, so per-user memory is what makes a purchased agent feel like a returning assistant rather than a stateless chatbot. The capability is load-bearing for the monetisation model. What is not documented is retention period, or how an operator or end user inspects, edits or deletes stored memories, beyond the fact that the API can list them. |
| Control & trust | ||
|
Human Oversight & Guardrails Approval steps, consent checkpoints, escalation rules, structured guardrails, policy constraints, and pause/resume controls. |
No / Not documented
P>N under the 31 August ruling that capability supplied by the vendor's people is not platform capability. This cell was re-based this morning onto the about page and has now lost that basis entirely, so it is re-decided on the product alone. WHAT THE RULING REMOVES. Every solution includes human oversight WHERE REQUIRED, with approval checkpoints and exception management, and the named solution shape of a multi-agent orchestration layer with approvals and escalation logic, both describe what the vendor's team architects into an engagement. A bespoke approval step built by the vendor for one client is not a mechanism the platform ships. WHAT IS LEFT ON THE PRODUCT, WHICH IS NOTHING. The API documentation enumerates the agent object's components as instructions, knowledge bases, forms and utilities. No approval, checkpoint, escalation, review, pause or pending-action object appears among them, and none appears in the Python SDK, the cookbooks or the starter repositories. Three passes have now searched approval and human-in-the-loop terms; the third returned only general guardrails literature about other vendors, which is a failed pass and is recorded as such rather than counted as confirmation. ONE NEAR-MISS REFUSED. Forms let an agent collect structured information from a person and trigger follow-up tasks from the response. A human supplying input mid-conversation is data capture, not oversight of agent action, and reading it as an approval gate would be the same error as crediting a governance essay. THE VENDOR'S OWN POSITIONING POINTS THE SAME WAY. A founder interview describes agents AUTONOMOUSLY APPROVING INVOICES WORTH CRORES through AI-driven validation. Where approval happens, the agent is the approver. N asserts the capability is not documented as first-class, which is what the evidence supports. It does not assert that engagements lack human checks; those are real and belong in the description rather than the grid. |
Partial
Stands at P, with the guardrail half strengthened by a mechanism the July basis did not have and the human half still absent. THE ADDITION IS SANDBOXED AGENT EXECUTION THROUGH THE OPENCLAW ENGINE, offered to regulated-industry customers. Execution isolation is a genuine containment guardrail and a stronger one than configuration limits, because it constrains what an agent can do at runtime rather than what it was told to do. It is named as an engine rather than a policy, which suggests a real boundary. AROUND IT SIT THE CONSTRAINTS THE JULY BASIS FOUND: access groups across public, member and invite-only, per-user usage budgets, role-based access controls and granular permissions, and per-file context instructions steering how documents are used. Those are real limits on who reaches an agent, how much they can consume, and what it draws on. WHAT IS STILL MISSING IS THE HUMAN HALF, and the July basis was right to say so. Nothing pauses an agent mid-run pending a person's approval, no pending-action queue exists, and no escalation or handoff path to a human is documented on any page reached across two passes. For a platform whose Actions write to HubSpot and Salesforce and whose agents are sold to third-party end users, the absence of an approval step before an external write is the notable gap. THE COMMERCIAL SHAPE EXPLAINS IT WITHOUT EXCUSING IT. Pickaxe agents are sold BY a creator TO end users, so the creator is not present during a run and there is nobody to approve; oversight is necessarily front-loaded into configuration and after-the-fact review. That is a coherent design and Partial describes it accurately. Prompt Injectors appear in the API documentation and may bear on runtime constraint; the page was not reached and is worth checking at lane close. |
|
Security, Identity & Governance RBAC, SSO, auditability, encryption, least-privilege tool access, compliance posture, and data handling policy. |
Partial |
Full / Explicit
Stands at F, and confidence rises because the July basis rested on one sentence from the features page while two dedicated security pages exist that it never reached. THE CONJUNCTION IS MET ON BOTH HALVES, WHICH IS WHAT THE 30 AUGUST BAR ASKS. Attestation: the vendor states it has been INDEPENDENTLY EXAMINED FOR SECURITY, AVAILABILITY, AND CONFIDENTIALITY CONTROLS, that SOC 2 REPORTS ARE AVAILABLE TO ENTERPRISE CUSTOMERS THROUGH OUR TRUST CENTER, and claims GDPR and CCPA compliance. Named customer-facing controls: enterprise single sign-on, role-based access controls, granular permissions, comprehensive audit logging, and configurable data retention. Four named controls where the bar requires one. A REPORT THAT EXISTS AND CAN BE REQUESTED IS MATERIALLY ABOVE THE HEDGE LADDER, whose rungs describe assertions with no report behind them. This is a completed examination with three trust services criteria named and a Trust Center as the delivery route. THE ONE HEDGE WORTH FLAGGING, AND IT IS A DELIBERATE WORD CHOICE: the vendor says EXAMINED throughout and never says TYPE II. Type is unspecified on every page reached, which on the ladder alone would sit at Partial. I have graded Full on the conjunction, since the controls half is documented in detail and a report is obtainable, but a procurement team should establish the type and period through the Trust Center before relying on it. That is the single thing to verify on this cell. TWO SUBSTANTIVE COMMITMENTS BEYOND THE CHECKLIST: knowledge base content and conversation data are NEVER USED TO TRAIN AI MODELS, and regulated-industry customers are offered sandboxed agent execution through the OpenClaw engine. Per standing practice the deployment position is graded on Dep, where it is None, and is not borrowed here. |
|
Observability & Auditability Traces, logs, execution histories, metrics, audit events, and debugging detail for production agent behavior. |
No / Not documented
P>N under the 31 August ruling that ongoing operational support is not observability. I dropped this to low confidence this morning while holding Partial, which was half the correction; the ruling completes it. WHAT THE RULING REMOVES. Ongoing operational support to MONITOR PERFORMANCE, REFINE WORKFLOWS, IMPROVE ACCURACY describes the vendor's staff watching the agents. I held Partial partly on the reported thirteen million agent actions in production, reasoning that a countable figure implies instrumentation. That reasoning was wrong in the same way customer logos implying security maturity was wrong on mobagel: the vendor having telemetry says nothing about whether the customer can see it. WHAT IS LEFT ON THE PRODUCT, WHICH IS NOTHING. No trace, run history, execution log, reasoning record, audit trail, dashboard or run-inspection endpoint appears in the API documentation, the Python SDK, the cookbooks or the starter repositories. The documentation covers creating agents, feeding knowledge bases, collecting form responses and invoking utilities, and stops there. THE ASYMMETRY THIS LEAVES IS THE MOST USEFUL THING ABOUT THIS RECORD. The platform has a real, published surface for BUILDING agents and no surface at all for WATCHING them. A customer can create an agent from code and cannot ask what it did afterwards. WHY THAT MATTERS MORE HERE THAN IT WOULD ELSEWHERE: the vendor's own pitch includes autonomously approving invoices worth crores and reconciling financial data. Those are exactly the actions someone is later asked to justify, and the party who can reconstruct them is the vendor rather than the buyer. N asserts the capability is not documented as first-class, which the evidence supports. Monitoring plainly occurs; it is the vendor's, and per the ruling that is not the platform's. |
Full / Explicit
P>F. The July basis withheld Full because detailed run tracing or audit logging was NOT DOCUMENTED AS FIRST CLASS. It is documented as first class on two pages the build did not reach. THE AUDIT HALF IS EXPLICIT AND COMPLIANCE-FACING: COMPREHENSIVE AUDIT LOGGING TRACKS EVERY INTERACTION, GIVING YOUR COMPLIANCE TEAM THE DOCUMENTATION THEY NEED FOR PROCUREMENT REVIEWS AND REGULATORY AUDITS, with the security page adding that operators track every interaction, monitor usage patterns and know exactly how the AI is being used, who is using it and when. THE INSPECTION HALF IS THE PART THAT DECIDES IT, and it comes from the user manual. An ACTIVITY tab shows usage across agents, EACH CONVERSATION HAS A HIGH-LEVEL AI-GENERATED SUMMARY, AND YOU CAN CLICK IN TO SEE THE FULL CONVERSATION IF SOMETHING NEEDS INVESTIGATING. Summary-then-drill-in is a genuine investigation workflow rather than a log dump: an operator with hundreds of conversations can find the one that went wrong, which is what auditability means in practice. FOR A CONVERSATIONAL AGENT THE TRANSCRIPT CARRIES MOST OF THE WHY. Unlike a background automation where reasoning is invisible unless traced, an agent that talks leaves its reasoning in the conversation, so full-transcript retrieval plus per-user attribution reconstructs what happened and largely why. THE HONEST GAP, and the reason confidence is medium: Action execution is not separately recorded on any page reached. For a platform with 500-plus Actions writing to HubSpot and Salesforce, a per-Action call record with parameters and results is the missing artefact, and a compliance team could see that a conversation occurred without seeing which CRM record it changed. Usage tracking, per-user credit accounting and analytics sit alongside. Retention period and export are undocumented. |
|
Deployment & Data Residency Deployment modes and options, including SaaS, dedicated cloud, VPC, on-prem, hybrid, local runtime, and self-hosting. |
Partial
F>P under the 31 August ruling that a bespoke deployment into a client environment is not a documented deployment option. I graded this Full this morning and the ruling is right that the basis did not support it. WHERE MY REASONING WENT WRONG. I treated the SDK's self-hosted domain parameter as corroborating the about page's client-owned-environment claim, and concluded the pair was stronger than either alone. But the two facts are not the same claim. The about page describes the vendor deploying a built solution into a client's environment as part of an engagement, which the ruling excludes. The SDK parameter proves only that the client library can be pointed at a non-default host. WHAT SURVIVES, AND IT IS GENUINELY SOMETHING. The SDK documents a domain parameter as required ONLY FOR SELF-HOSTED AI LIBRARY INSTANCES. That is a first-party, code-level artefact stating self-hosted instances exist, and it serves no marketing purpose, which is why it is worth more than the sales sentence. Customer-hosted deployment is therefore real rather than asserted. WHY THAT IS PARTIAL AND NOT FULL. Nothing documents how a customer obtains or runs one. There is no installation guide, no infrastructure requirement, no statement of which components run customer-side, no region selection, and no indication whether self-hosting is a purchasable option or something the vendor stands up during a delivery engagement. A capability that exists but has no documented route for a customer to take is exactly the restricted case the ground rules place at Partial, with the limitation named. THE GENERAL LESSON WORTH CARRYING, since it will recur: a code artefact proves a capability exists; it does not prove the capability is offered. Those are different questions and this axis asks the second. The managed option is the documented default and is not in question. |
No / Not documented
P>N, and the correction is a category one: the July basis graded distribution surfaces as deployment. WHAT IT CITED was rich multi-surface deployment through website embeds, branded portals with custom domains, API, email, WhatsApp and Slack. Every one of those is a channel by which an end user REACHES the agent, and every one is already credited on Trig, where they belong and where they earn Full. Where a chat widget renders says nothing about where the software runs or where the data sits, which is what this axis measures. The same basis then conceded the actual answer: THE PLATFORM IS CLOUD ONLY AND NO SELF HOSTED OR DATA RESIDENCY OPTIONS ARE DOCUMENTED. THIS PASS CONFIRMS THAT CONCESSION AGAINST THE VENDOR'S OWN SECURITY PAGES, which are where a residency option would be advertised if one existed. They are unusually detailed, covering encryption at rest and in transit, data ownership, retention configuration, a Trust Center and custom SLAs, and they name no region selection, no virtual private cloud, no on-premises option and no data localisation commitment. A vendor documenting this much and omitting residency is not hiding it. UNDER SECTION 7 SINGLE-REGION SAAS WITH NO SELECTION IS NONE, and that is the correct reading. Grading Partial here would credit the platform twice for the same channel list. TWO ADJACENT FACTS RECORDED AND DELIBERATELY NOT CREDITED. Retention configuration is a control and is credited on Sec. SANDBOXED AGENT EXECUTION THROUGH THE OPENCLAW ENGINE, offered to regulated industries, is runtime isolation rather than deployment location, and is credited as a guardrail on HITL. ADVANCED DATA PRIVACY OPTIONS are offered to enterprise customers without being specified. If those turn out to include residency, this cell moves; it is the natural question for a lane-close pass through the Trust Center. |
| Solution readiness | ||
|
Prebuilt Agents, Templates & Packs Ready-made workflows, packaged employees, templates, blueprints, industry solutions, and role-specific agents that reduce time-to-value. |
Partial
Stands at P, with better evidence than the reusable-components claim the July basis used. WHAT IS ACTUALLY PUBLISHED AND ADOPTABLE. The vendor's GitHub organisation carries a COOKBOOKS repository and starter application repositories including chat agent implementations in Next.js and Angular 19 and an event template. Those are real artefacts a customer takes and runs, which is more than the July basis could point at, and under the 31 August bar packaged assets ready to adopt by any route is the test. SOLUTION ARCHETYPES ARE NAMED rather than gestured at. The vendor describes architecting a conversational agent, document workflow, research assistant, email automation system, or a multi-agent orchestration layer with approvals and escalation, and covers use cases across sales, service, finance, operations and document processing. WHAT HOLDS IT BELOW FULL is what those artefacts are. Starter apps are UI scaffolding for embedding an agent, and cookbooks are worked examples; neither is a packaged agent a customer adopts to do a job. The archetypes are descriptions of what the vendor builds for clients, not a catalogue a customer browses. No template gallery, agent marketplace or prebuilt agent library appears on any surface reached. THE DISTINCTION THAT DECIDES IT, and it recurs across this record: the reusable components are the vendor's own delivery accelerators, used by AI Library Code when the vendor assembles a solution. The customer receives the finished solution, not the components. Assets that speed the builder are not assets shipped to the buyer, and only the second is what this axis measures. |
Partial
Stands at P, with one fact removed for working two axes and one new lead recorded that could move it. WHAT CAME OUT: the prebuilt Actions library. Actions are integrations and are credited on Int, where they carry Full on an enumerated public directory. Counting the same library here would credit one catalogue twice, and it is not what this axis measures anyway: a connector is a capability the customer wires in, not a packaged agent they adopt. WHAT REMAINS IS GENUINELY MIXED. A gallery of use-case examples shows what others have built, which is illustration rather than an adoptable asset. The AUTOMATIC AI BUILDER generates an agent from a description, which is the kalcend pattern seen earlier today: generation substituting for a catalogue. Neither is a packaged asset ready to adopt. THE LEAD THAT COULD MOVE THIS CELL, recorded rather than graded: site furniture on the Actions pages reads GET STARTED FASTER WITH PRE-BUILT TEMPLATES, CHOOSE FROM OUR LIBRARY OF READY-TO-USE AI TOOLS AND CUSTOMIZE THEM FOR YOUR NEEDS. Under the 31 August bar a library of ready-to-use tools would clear Full, and the phrasing is a direct first-party claim. I have not graded from it because I reached the sentence and not the library, and I am not willing to grade a catalogue I have not seen after removing another catalogue for being counted twice. Locating that template library is the single check for this cell at lane close. WORTH NOTING FOR THE COMMERCIAL MODEL: Pickaxe's premise is that creators PACKAGE AND SELL agents, so the marketplace it operates is the creators' storefronts rather than a vendor-supplied pack library. The absence is coherent with the business rather than an oversight. |
| Platform extensibility | ||
|
Model Flexibility & Routing Ability to work across multiple foundation models, route tasks to different models, or let buyers bring their own providers and keys. |
No / Not documented
Stands at N, and the July basis was right. Confidence rises from low to medium-strength reasoning even though the numeric stays low, because the absence is now observed on a first-party surface rather than merely unfound. WHAT THE DOCUMENTATION ACTUALLY SAYS. The API documentation enumerates the components of an agent, instructions, knowledge bases, forms and utilities, and states that agents BY DEFAULT PICK ON THE LLM'S KNOWLEDGE. No model, provider, or routing parameter appears among the agent's documented properties, and no provider is named anywhere on the documentation index, the product page or the SDK repository. THIS IS THE WEAKER FORM OF ABSENCE, and the note should say so. Per the ground rules, a vendor naming a single provider is the best-evidenced absence on this axis, because the absence is visible rather than inferred. Here nothing is named at all, so the reader learns only that the platform does not present model choice as a feature, not which model powers it. ONE ADJACENT FACT RECORDED AND NOT CREDITED: a founder interview describes the launch of the OpenAI APIs as the turning point that made the company possible. That is origin narrative, not a disclosure of what powers the product today, and it certainly is not customer selection. CONFIDENCE STAYS LOW DELIBERATELY. I reached the documentation index but not a full agent API reference, and a model parameter on the agent creation endpoint is exactly the kind of thing that lives one level below an overview page. This is the cheapest cell on the record to overturn and it should be checked before this grade is relied on. |
Full / Explicit
Stands at F and is one of the cleanest Model cells in the lane, because the customer's choice is both real and cheap to revise. MORE THAN 50 MODELS ACROSS OPENAI AND ANTHROPIC with ONE-CLICK SWITCHING. The switching cost is what earns Full rather than the count: a choice made at build time and then frozen is weaker than one a creator revisits when a cheaper or better model ships, and one click is as low as that cost goes. THE COST DIMENSION IS THE PART WORTH CARRYING and it is unusual in this lane. Live model rates are shown in the product, and a cost estimator lets a creator compare price against speed and quality before committing. Pickaxe's economics make this load-bearing rather than decorative: creators resell agents and end-user usage draws down the creator's credits at one dollar of credit per dollar of model cost, so model choice is directly the creator's margin. A platform that shows live rates at the point of selection is helping its customers price their own products. TWO PROVIDERS RATHER THAN A LONG PROVIDER LIST is the honest limit. Google, Meta, Mistral and open-weight models are not documented, so a customer wanting a specific non-OpenAI, non-Anthropic model cannot have it, and there is no bring-your-own-key or custom endpoint path documented. Within those two families the depth is real. No automatic routing between models at runtime is documented, which this axis does not require; it measures customer choice, which is amply met. Per section 7 the Actions surface and the API run the other direction and are credited on Int and Ext rather than counted here. |
|
APIs, SDKs & MCP Extensibility Composability layer: stable APIs, SDKs, MCP tool consumption/serving, custom tools, and integration into internal systems. |
Full / Explicit
P>F, and this is the correction the record needed. The July basis said a public developer API, SDK or MCP endpoint WAS NOT DOCUMENTED THIS SESSION. All three exist and are published under the vendor's own name. WHAT THE PASS FOUND. A public Python SDK at github.com/ailibrarycloud/ailibrary-python, described as a framework for building agents and multi-agent orchestrations, authenticating with an AI_LIBRARY_KEY against a REST API at api.ailibrary.ai. The documented calls are real platform operations, not a wrapper: client.agent.create with title and instructions, and client.files.upload binding uploaded documents to an agent's knowledge_id. A full API documentation site sits at docs.ailibrary.ai, and the vendor publishes a documentation repository, a cookbooks repository and starter application repositories in the same organisation. THE DETAIL WORTH CARRYING is a single comment in the SDK: the domain parameter is ONLY REQUIRED FOR SELF-HOSTED AI LIBRARY INSTANCES. That one line independently corroborates the client-owned deployment claim graded on Dep, which otherwise rested on a marketing page. AI LIBRARY MCP was announced around May 2026 as a unified infrastructure layer acting as a single server giving coding agents structured access to tools, data and workflows across search, parsing, storage, retrieval and APIs. That is the MCP direction this axis credits, an outside agent reaching in. It is carried by vendor announcement quoting the founders rather than a documentation page, so it corroborates rather than carries the grade; the SDK and API alone clear Mike's 30 August bar. WHY THE JULY PASS MISSED IT is worth recording for the lane: the build worked from the homepage, a second-domain about page, LinkedIn and Crunchbase, and never searched the vendor's name against GitHub or a docs subdomain. For a small vendor that is where the developer surface lives. |
Full / Explicit
Stands at F, and it is now among the better-evidenced Ext cells in the lane because a real API documentation site exists that the July build did not cite. THE DOCUMENTATION IS STRUCTURED AND SPECIFIC, not a contact-sales page. pickaxe.co/v1/documentation covers an Introduction, a COMPLETIONS API for deployment inference requests, CLI and Coding Agents, Prompt Injectors, Environment Variables, Webhooks and Authenticated Embeds, plus RESOURCE-SPECIFIC CRUD AND LISTING ENDPOINTS with request and response examples. CRUD across resources means the platform is administrable from outside, not merely queryable, which is the stronger form of what Mike's 30 August bar asks. THE CLI IS THE DISTINCTIVE PART AND IT IS AIMED SOMEWHERE UNUSUAL. It is described as BUILT FOR AI CODING AGENTS LIKE CLAUDE CODE, CURSOR, AND WINDSURF, letting a coding agent create agents, sync knowledge bases and deploy WITHOUT OPENING THE WEB APP. A no-code platform shipping a machine-first control surface is a deliberate second audience, and it is the Ext direction precisely: the customer's assistant operates Pickaxe. WEBHOOKS RUN OUTBOUND to Zapier, Make and other endpoints, and Authenticated Embeds cover programmatic session handoff, so the surface spans inbound control, outbound events and embedded delivery. ONE FACT RECORDED AND NOT CREDITED: a community-built MCP server exists on GitHub exposing agents, knowledge bases, users and analytics to Claude. It is third-party rather than the vendor's, so it carries nothing, but its existence corroborates that the public API is complete enough for someone outside the company to wrap it. No first-party MCP server was found, which under the ruling does not withhold the grade. |
|
Testing, Debugging & Optimization Testing, debugging, scoring, retries, fallbacks, quality gates, and optimization loops for improving agent workflows before and after deployment. |
No / Not documented
P>N under the 31 August ruling that a delivery lifecycle with a testing stage is not an eval surface. WHAT THE RULING REMOVES, AND IT WAS THE WHOLE BASIS. Testing and optimisation as lifecycle stages, ongoing accuracy improvement after launch, and higher accuracy through deterministic logic and controlled orchestration all describe how the vendor builds and tunes solutions. The vendor tests; the customer does not. I HELD PARTIAL THIS MORNING PARTLY ON OUTCOME-BASED PRICING, reasoning that billing on invoices reconciled and tickets resolved implies measurement rigorous enough to settle an invoice. The ruling is right that this does not survive. A commercial settlement figure is a business outcome agreed between two parties, not a result the customer can read about the agent's behaviour and compare across versions. If anything the pricing model points the other way: outcome pricing exists precisely so the buyer does not have to evaluate the system, only the result. WHAT IS LEFT ON THE PRODUCT, WHICH IS NOTHING. No test set, expected outputs, scored run, judge, harness, regression comparison or version comparison appears in the API documentation, the Python SDK, the cookbooks or the starter repositories. THE RULED DEFINITION OF PARTIAL DOES NOT FIT EITHER, and this is worth stating because I applied it wrongly this morning. Partial is a quality gate with no readable result. A quality gate is still a mechanism the customer's work passes through. Here there is no gate in the product at all; there is a supplier who tests before handing over. That is None. RECORDED FOR CONSISTENCY: this is the same distinction that keeps mobagel at Partial rather than Full, real measurement of the wrong object, applied one step further along, to real measurement by the wrong party. |
Partial
Stands at P under the 31 August Eval bar, and the reasoning is worth recording because this record has more testing surface than most Partials. WHAT EXISTS IS A REAL PRE-DEPLOYMENT LOOP. Preview and test mode, IMPERSONATING SPECIFIC USERS to test access paths, one-click model switching for comparison, a cost estimator, and AI-POWERED PROMPT IMPROVEMENTS WITH REAL-TIME TESTING. User impersonation is the standout: for a platform whose agents sit behind access groups and paywalls, testing what a member sees versus an invite-only user is exactly the failure mode that would otherwise reach a paying customer. THE OPTIMISATION LIMB IS GENUINELY SERVED. Comparing models on cost, speed and quality with live rates in view is optimisation with a readable number attached, and it is more than most vendors at this grade offer. WHY IT STOPS SHORT OF FULL. The ruled bar asks for a readable comparable result about THE AGENT'S BEHAVIOUR on the customer's own work. Cost and latency are properties of the model, not judgements about whether the agent answered correctly. Nothing scores an output, retains a test set, records expected answers, or lets a creator establish whether a prompt or knowledge change improved behaviour rather than merely changing it. Real-time testing shows what happens; it does not tell you whether it was better. THE GAP IS SHARPER HERE THAN THE GRADE CONVEYS, and worth carrying. Creators sell these agents to paying end users and the knowledge base re-syncs daily, so behaviour drifts without anyone changing anything. A platform with automatic content refresh and no regression check is one where quality can degrade silently between deployments. The Activity tab's conversation summaries support after-the-fact review and are credited on Obs. |
| Specialist automation | ||
|
Browser & Computer Use Browser, desktop, or remote/local computer control for workflows that cannot be handled through stable APIs alone. |
No / Not documented
Stands at N, and confidence rises from low to high because a near-miss was found and refused rather than nothing being found at all. That is a materially stronger basis than the July one, which rested on having seen no claim. THE NEAR-MISS, AND IT IS THE MOST COMMONLY MISCOUNTED ONE IN THE INDEX: the utilities set lets an agent SCRAPE A WEB PAGE. A vendor whose agent fetches and parses web pages will be described by third parties as browsing the web, and a grader working from a summary could easily read that as computer use. IT IS NOT. Scraping retrieves a page's content over HTTP and parses the markup. Comp is non-zero only where an agent operates software through its interface BECAUSE no programmatic interface exists. Fetching a URL is the programmatic interface. Nothing navigates, clicks, fills a field or reads a rendered screen. The same refusal applies to the adjacent utilities, web search, news search, and parsing of PDFs and handwritten text. Handwriting parsing is optical character recognition on a document, which is document intake rather than interface operation, and it is credited on Know. Everything else the platform does runs through documented programmatic paths: the REST API, the SDK, knowledge bases reading databases, and forms collecting structured input. This is consistent with the scraping, screenshot, sandbox and browser-testing refusals made repeatedly across this lane, and against sema4-ai at Full on genuine desktop automation. The line has been applied the same way each time. |
No / Not documented
Stands at N, and confidence rises to high because this pass found two near-misses and refused both, which is a stronger basis than the July one of having seen no claim. NEAR-MISS ONE, ALREADY REFUSED IN JULY AND CORRECTLY: a Google search Action gives real-time web retrieval. Fetching results is retrieval, not operating an interface, and the July basis said so. NEAR-MISS TWO IS NEW AND IS THE ONE THAT WOULD CATCH A GRADER SKIMMING. The user manual lists CODE EXECUTION, RUN CODE DIRECTLY FROM THE AGENT, alongside PDF generation. Code execution is the most frequently miscounted fact on this axis and it was the exact defect corrected across thirteen records in the Coding agent June cohort. Running code is not operating software that lacks a programmatic interface; it is using the most programmatic interface there is. EVERYTHING ELSE RUNS THROUGH DOCUMENTED PROGRAMMATIC PATHS: more than 500 Actions and MCP servers, webhooks, the Completions API, and a vector knowledge base. The platform's entire premise is that a non-technical creator connects tools that already expose APIs. THE ABSENCE IS STRUCTURAL RATHER THAN A GAP, and worth stating so the cell is not misread as a deficiency. Pickaxe agents are conversational products embedded in a page, a portal, Slack or WhatsApp, answering a person. Driving a third-party interface has no place in that shape, and no page reached makes any claim in that territory. Recorded for the lane: this is the third record this session where code execution or scraping presented as computer use and was refused, against kalcend, where genuine browser navigation with an OTP handoff earned Full. The line has been applied identically in both directions. |
Pricing snapshot
Sourced from the Index pricing dataset · open each vendor's profile for full detail.
| Pricing | ||
|---|---|---|
|
Entry price Lowest public entry point |
Free to start; enterprise delivery priced by sales | Free Starter plan; Gold at twenty nine dollars per month and Pro at ninety seven dollars per month, plus credit based usage billed at cost. |
|
Pricing confidence How public the numbers are |
Public, partial | Public, exact |
|
Billing Primary billing axis |
— | Monthly plan plus credit based usage at underlying model cost |
|
Variable cost Workload / overage exposure |
Medium variable cost | High variable cost |
|
Free tier / trial Try before you buy |
Free tier
|
Free tier
|
|
Buying motion Self-serve vs sales call |
Mixed | Self-serve |
More comparisons with AI Library or Pickaxe
Other matchups in agent builders
Not the pairing you were after? These compare a different set of agent builders on the same 14 capabilities.