Claude API Integration: Add Claude to Your App
How to integrate Anthropic's Claude API into existing software: the Messages API flow, JSON outputs, cost per request from Anthropic's October 2026 prices, rate-limit tiers and data terms.
- Organisation
- WavX Solutions
- Telephone
- +919310079927
Description
Integration Anthropic Claude API integration: adding Claude to your app or system
A Claude API integration calls Anthropic's models from your server through the Messages API, asks for output in a JSON schema, and checks it before anything is saved. As of October 2026 Anthropic lists Claude Haiku 4.5 at $1 per million input tokens, Sonnet 5.5 at $2 and Opus 5.5 at $4, with cache reads and batch jobs charged far less, so prompt design sets the running cost as much as model choice.
Discuss an integration Last updated 2 October 2026
What the integration connects
Anthropic's API gives your software access to the Claude models. The integration is the service in between: it takes a record or document from your system, sends it to Claude with fixed instructions, receives a structured answer, checks it, and writes it back where your staff work.
Good first uses are narrow and frequent. Reading supplier invoices into purchase entries, tagging support tickets, summarising long customer threads for the account manager, or answering staff questions from a policy manual. Where the work must take actions in other systems, it becomes an AI workflow automation project; where it only adds a feature to an existing screen, it is covered by add AI to your existing software .
Data flow: reading a supplier invoice into a purchase entry
Accounts team uploads invoice PDF in your app
Your server ── fixed system prompt (company rules, ledger list) ── marked for caching
── the PDF as a document block
── output_config.format: json_schema {supplier_gstin, invoice_no,
date, lines[], taxable_value, cgst, sgst, igst, total}
POST https://api.anthropic.com/v1/messages (model: claude-sonnet-5-5)
Response ── stop_reason: end_turn? max_tokens? refusal?
── parse JSON; re-check totals and GSTIN format in code
── usage: input, cache read, output tokens → cost logged
├── totals match ──▶ draft purchase entry, flagged "AI-read, check"
└── mismatch ──────▶ manual queue with the PDF and the AI's reading
The arithmetic is checked by your code, not trusted from the model. A schema makes sure the fields exist and have the right types; it says nothing about whether the numbers add up.
Field mapping
Your system
Messages API field
Notes
Company rules, ledger names, output instructions
system
Stable text; mark it for prompt caching so repeats are billed at the cache-read rate
The document or record
messages (user turn: text or document block)
Only what the task needs
Output shape
output_config.format with type: json_schema
Objects need additionalProperties: false
Model
model
claude-opus-5-5 , claude-sonnet-5-5 or claude-haiku-4-5 ; kept in configuration
Length cap
max_tokens
If reached, stop_reason is max_tokens and the JSON may be incomplete
Outcome
stop_reason
Handle refusal and max_tokens before parsing
Cost tracking
usage (input, cache creation, cache read, output tokens)
Stored per request
Data location
inference_geo ( global or us )
us costs 1.1 times the standard rate
Cost per request
Prices are Anthropic's per million tokens, read on 2 October 2026, in US dollars.
Model (Anthropic's description)
Input
Cache read
Output
One request: 3,000 in, 400 out
10,000 such requests
Claude Haiku 4.5 ("the fastest model with near-frontier intelligence")
$1
$0.10
$5
$0.005
$50
Claude Sonnet 5.5 ("the best combination of speed and intelligence")
$2
$0.20
$10
$0.010
$100
Claude Opus 5.5 ("long-running agentic coding and knowledge work")
$4
$20
$0.020
$200
The calculation is tokens × price ÷ 1,000,000, input and output added, with no caching. Three levers change it:
Prompt caching. A cache read costs 10% of the input price on most models and 5% on Opus 5.5. Writing to the 5-minute cache costs 1.25 times input and to the 1-hour cache 2 times. A long fixed system prompt repeated on every invoice is where this pays.
Batch processing. The Batch API is 50% off input and output, for work that can wait, such as last month's invoices.
The tokenizer. Anthropic says Claude 4.7 and later models produce about 30% more tokens for the same text than earlier ones. Estimate from a real count, not a word count.
Limits and gotchas
Issue
What Anthropic documents
What the integration does
Rate limits
Requests, input tokens and output tokens per minute, per model; at the Start tier, 1,000 requests per minute for Opus 5.5, Sonnet 5.5 and Haiku 4.5
Queue; obey the retry-after header on 429
Cached input
For most models, cache reads do not count towards the input-token rate limit
Cache the stable prefix; it raises throughput as well as cutting cost
Monthly spend caps
Start $500, Build $1,000, Scale $200,000; at the cap, requests return 429 until the next month
Alert well before the cap; request a higher tier ahead of launch
New organisations
May begin on an evaluation tier below the standard limits
Do not plan a launch on day-one limits
Schema features
No recursive schemas, no numeric or string-length constraints; enum capitalisation is not guaranteed
Enforce ranges and compare enums case-insensitively in code
Model retirement
Haiku 4.5's retirement is listed as not sooner than 15 October 2026; Opus 5.5 not sooner than 22 September 2027
Keep a test set so a model change is a re-run, not a rewrite
Inference: global or US only; workspace storage: US only
Record this in your privacy notice; strip data the task does not need
What WavX builds
WavX builds the service around Claude: input preparation and redaction, the system prompt and schema under version control with caching in mind, validation in code of every answer, retries and queueing within the rate limits, per-request cost logging, batch jobs for back-office work, and a review screen for staff. The same service can call OpenAI instead; see OpenAI API integration and the OpenAI, Claude and Gemini API comparison.
Effort band
The AI cost calculator prices an AI feature in an existing app from a ₹2,50,000 base: a planning range of ₹2,12,500 to ₹3,12,500 with no custom data, rising when the feature works on your documents (×1.25) or live systems (×1.5). These are planning ranges, not quotes. Usage is billed by Anthropic to your own account.
When not to integrate the API
If the job is occasional drafting or summarising by a few staff, Claude's own apps cost less than building anything. If the task follows fixed rules (GST rate by HSN code, routing by pincode), write ordinary code. If invoices arrive as structured e-invoice JSON already, read the JSON; a model adds cost and risk to data that is already machine-readable.
Frequently asked questions
Should we use Claude or OpenAI?
Test both on a few hundred of your own real cases; the difference for a given task shows up there, not in benchmarks. Both are priced per token and both offer structured JSON output. We build the service so the model is a configuration setting, which keeps the choice open and lets you move if prices or quality change.
Does Anthropic train on what we send through the API?
Anthropic's privacy centre says that by default it does not use inputs or outputs from commercial products, including the API, to train its models. It also says API inputs and outputs are deleted from its back end within 30 days, with exceptions such as the Files API, custom agreements, policy enforcement and legal requirements.
Can inference run in India?
Not as of October 2026. Anthropic's data residency page lists two inference settings, global and US-only, and US as the only workspace storage location. US-only inference costs 1.1 times the standard rate on Claude 4.6 and later models.
Which model should we start with?
Anthropic's own guidance is to start with Claude Opus 5.5 for most workloads. For a high-volume, narrow task such as extraction or tagging, we test Sonnet 5.5 and Haiku 4.5 against the same cases and pick the cheapest one that passes.
Sources
Claude docs: Pricing · read 2 October 2026
Claude docs: Models overview · read 2 October 2026
Claude docs: Rate limits · read 2 October 2026
Claude docs: Structured outputs · read 2 October 2026
Claude docs: Data residency · read 2 October 2026
Anthropic Privacy Center: How long do you store my organisation's data? · read 2 October 2026
Anthropic Privacy Center: Is my data used for model training? · read 2 October 2026
Related
Add AI to your existing software
OpenAI API integration
AI cost calculator
Build your own software — your way, your pricing.
WavX Solutions is here to create your own software in a fully custom way, built exactly how you work — with a pricing model that fits your business. Connect now and let's build it.
Contact Now helpwavx@gmail.com