Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API · Core API calls across providers · lesson 5 of 19 · 17 min
Tool calling across Claude, OpenAI and Gemini
Same idea, three dialects
Tool (function) calling lets a model request that your code run a function with JSON arguments. The loop is identical everywhere: send tools → model returns a call → you execute → you send the result → model continues. The shapes differ:
| Step | Claude Messages | OpenAI Responses | Gemini | |---|---|---|---| | Define a tool | {"name", "description", "input_schema": {...}} (optional "strict": true) | {"type": "function", "name", "description", "parameters": {...}, "strict": true} | types.FunctionDeclaration(name, description, parameters_json_schema={...}) inside types.Tool(function_declarations=[...]), or pass Python functions directly | | Model asks | stop_reason == "tool_use"; tool_use blocks with id, name, input (dict) | output items with type == "function_call", name, arguments (JSON string), call_id | response.function_calls with name, args (dict) | | You reply | user message with tool_result blocks (tool_use_id, content, is_error) | input item {"type": "function_call_output", "call_id", "output"} | types.Part.from_function_response(name=..., response={...}) in a role="tool" content | | Force/choose | tool_choice (auto, any, tool, none; some newest models only allow auto/none) | tool_choice | tool_config function-calling mode (AUTO, ANY, NONE) |
Gemini's Python SDK can also run automatic function calling when you pass Python callables as tools; OpenAI's and Anthropic's SDKs offer tool-runner helpers (Anthropic's is beta). Helpers are convenient, but you still need to understand the manual loop for approvals, logging and error handling.
One tool, three implementations
The tool: get_order_status(order_id) for a Karachi e-commerce store.
import json, os
import anthropic
from openai import OpenAI
from google import genai
from google.genai import types
ORDERS = {"KHI-1042": {"status": "out_for_delivery", "eta": "today 6-9pm"}}
def get_order_status(order_id: str) -> dict:
return ORDERS.get(order_id, {"error": f"Order {order_id} not found; ask the customer to check the ID"})
SCHEMA = {"type": "object", "properties": {"order_id": {"type": "string", "description": "Like KHI-1042"}},
"required": ["order_id"], "additionalProperties": False}
DESC = "Look up delivery status and ETA for one order ID."
Q = "Where is my order KHI-1042?"
# ---- Claude ----
cl = anthropic.Anthropic()
tools = [{"name": "get_order_status", "description": DESC, "input_schema": SCHEMA, "strict": True}]
msgs = [{"role": "user", "content": Q}]
r = cl.messages.create(model=os.environ.get("CLAUDE_MODEL", "claude-sonnet-5"), max_tokens=1000, tools=tools, messages=msgs)
while r.stop_reason == "tool_use":
msgs.append({"role": "assistant", "content": r.content})
results = [{"type": "tool_result", "tool_use_id": b.id, "content": json.dumps(get_order_status(**b.input))}
for b in r.content if b.type == "tool_use"]
msgs.append({"role": "user", "content": results})
r = cl.messages.create(model=os.environ.get("CLAUDE_MODEL", "claude-sonnet-5"), max_tokens=1000, tools=tools, messages=msgs)
print("Claude:", "".join(b.text for b in r.content if b.type == "text"))
# ---- OpenAI Responses ----
oa = OpenAI()
otools = [{"type": "function", "name": "get_order_status", "description": DESC, "parameters": SCHEMA, "strict": True}]
items = [{"role": "user", "content": Q}]
resp = oa.responses.create(model=os.environ.get("OPENAI_MODEL", "gpt-5.5"), input=items, tools=otools)
while any(o.type == "function_call" for o in resp.output):
items += resp.output # keep the model's call items in the history
for o in resp.output:
if o.type == "function_call":
out = get_order_status(**json.loads(o.arguments)) # arguments is a JSON string
items.append({"type": "function_call_output", "call_id": o.call_id, "output": json.dumps(out)})
resp = oa.responses.create(model=os.environ.get("OPENAI_MODEL", "gpt-5.5"), input=items, tools=otools)
print("OpenAI:", resp.output_text)
# ---- Gemini (manual declaration) ----
gm = genai.Client()
decl = types.FunctionDeclaration(name="get_order_status", description=DESC, parameters_json_schema=SCHEMA)
cfg = types.GenerateContentConfig(tools=[types.Tool(function_declarations=[decl])],
automatic_function_calling=types.AutomaticFunctionCallingConfig(disable=True))
contents = [types.Content(role="user", parts=[types.Part.from_text(text=Q)])]
g = gm.models.generate_content(model=os.environ.get("GEMINI_MODEL", "gemini-flash-latest"), contents=contents, config=cfg)
while g.function_calls:
contents.append(g.candidates[0].content)
parts = [types.Part.from_function_response(name=fc.name, response=get_order_status(**fc.args))
for fc in g.function_calls]
contents.append(types.Content(role="tool", parts=parts))
g = gm.models.generate_content(model=os.environ.get("GEMINI_MODEL", "gemini-flash-latest"), contents=contents, config=cfg)
print("Gemini:", g.text)
Add a step cap to every while loop in production (these are kept short for comparison).
Cross-provider design rules
- Write one canonical JSON Schema per tool and generate each provider's format from it. Keep schemas simple: objects, strings, numbers, booleans, enums, arrays; avoid exotic keywords that one provider may not support in strict mode.
- Strict modes (Claude
strict: true, OpenAIstrict: true) guarantee arguments match the schema but requireadditionalProperties: falseand listing required fields. - Parse defensively: OpenAI returns arguments as a JSON string; Claude and Gemini return parsed objects. Never string-match raw JSON.
- Parallel calls: all three may request several tools at once. Execute them (concurrently if safe) and return all results together.
- Errors: return them as results the model can read (
is_erroron Claude; a structured{"error": ...}payload elsewhere). - Security: validate arguments and permissions in code; the model's request is a suggestion.
Server-side and hosted tools
Providers also offer tools that run on their infrastructure: web search, code execution, file search, remote MCP connectors and more. They save you building infrastructure but vary by provider, model and platform (for example, some are not available on every cloud platform). Check each provider's current tool list and data-handling terms before using them with sensitive data.
Pitfalls
- Forgetting to append the model's tool-call turn before the results (the API then rejects or misinterprets the results).
- Returning results out of order or dropping one of several parallel calls.
- Loops without step caps.
- Executing write actions without validation or approval.
Measuring success
Tool-selection accuracy and argument validity per provider on the same eval set, calls per task, and tool error rates. Differences between providers here are often larger than differences in plain text quality.
Video lecture: Tool calling across Claude, OpenAI and Gemini
Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.
- Tool calling across providers
- Why it matters
- The universal loop
- Claude dialect
- OpenAI and Gemini dialects
- Simple example: order status
- Business example: Dubai logistics
- Portable tool rules
- Testing tool calling
- Hosted tools
- Common mistakes
- How many tools?
- Deeper: the Dubai logistics monitor (illustrative)
- Watch me do it: order status loop
- Recap + try this now
Lecture transcript
Tool calling across providers
Tool calling is where language models stop being chatbots and start doing work: checking an order, booking a slot, updating a record. Every major provider supports it, but each speaks a slightly different dialect. In this lesson you'll learn the one loop that never changes, the three dialects of Claude, OpenAI and Gemini, and how to write tools once so they work with all of them.
Why it matters
Why does this matter? Because tools are how AI connects to your business: your orders, your calendar, your CRM. If you tie your tool code tightly to one provider, switching later means rewriting every integration. Think of an interpreter at a conference. The speaker's message is the same, but it's delivered in three languages. Your tool is the message. The provider format is just the language. Keep the message separate, and translation becomes easy.
The universal loop
Here's the loop that never changes. You send the model your request and a list of tools, each with a name, a description and a JSON schema for its arguments. The model either answers, or asks for a tool with specific arguments. Your code runs the function. You send the result back. The model continues, maybe calling another tool, until it answers. The model never runs anything itself. It proposes; your code decides and executes.
Claude dialect
Dialect one, Claude. You define tools with a name, description and input schema, and can set strict to true so arguments always match the schema. When Claude wants a tool, the stop reason is tool use, and the content has tool use blocks with an id, a name and the input as an object. You append the assistant's content, then send a user message containing tool result blocks, each pointing to its tool use id, with an is error flag if something failed.
OpenAI and Gemini dialects
Dialect two, OpenAI's Responses API. Tools have a type of function, a name, a description, parameters and a strict flag. The model's output contains function call items with a name, a call id, and arguments as a JSON string, so you must parse it. You add the model's items to your input list, then add function call output items with the matching call id and your result as a string. Dialect three, Gemini. You declare functions, wrap them in a tool object, and read response dot function calls, with arguments already parsed. You reply with function response parts in a tool role message.
Simple example: order status
A simple example first. One tool: get order status, for a Karachi online store. The customer asks, where is my order K H I ten forty two. All three providers call the tool with that order id, your code returns out for delivery, arriving today between six and nine in the evening, and each model writes a friendly answer. The lesson's code shows all three loops side by side. Read them together and you'll see the same four steps in three spellings.
Business example: Dubai logistics
Now a realistic business example. A logistics startup in Dubai supports customers in English and Arabic and wants the freedom to switch providers as prices change. They keep one canonical JSON schema per tool: track shipment, reschedule delivery and open a claim. A tiny converter generates each provider's format. The same eval of forty real customer messages runs against two providers every month. When one provider's tool selection accuracy dropped after a model update, they switched the routing in configuration and fixed nothing in their tool code.
Portable tool rules
A few design rules make tools portable. Keep schemas simple: objects, strings, numbers, booleans, enums and arrays. Strict modes need additional properties set to false and every required field listed. Parse arguments defensively, since OpenAI sends a string and the others send objects. When several tools are requested at once, run them and return all the results together. Send errors back as results the model can read. And always validate arguments and permissions in your code, because the model's request is only a suggestion.
Testing tool calling
How should you test tool calling across providers? Build a small eval of realistic user messages, each labeled with the tool you expect and the key arguments. For the logistics example, that might be: where's my parcel, with a tracking number, which should call track shipment; can you deliver on Saturday instead, which should call reschedule; and my box arrived crushed, which should open a claim. Run the same forty messages through each provider and score three things: right tool, valid arguments and calls per task. Differences here are often bigger than differences in writing quality.
Hosted tools
Providers also offer hosted tools that run on their infrastructure, like web search, code execution, file search and connectors to remote MCP servers. They save you building infrastructure, but availability varies by provider, model and cloud platform, and so do data handling terms. Use them deliberately, and check the current docs before sending sensitive data through them.
Common mistakes
Common mistakes. Forgetting to add the model's tool call turn to the history before sending results, which makes the API reject your request. Returning results out of order, or dropping one of several parallel calls. Loops with no step cap. And running write actions, like refunds, without validation or approval. Each of these shows up quickly in testing if you look for it.
How many tools?
Here's a question many teams ask: how many tools can I give a model before quality drops? There's no single number, and it varies by model and provider. The practical signal is your eval: if tool selection accuracy falls as you add tools, or calls per task rise, you've gone too far. Group tools into focused sets per feature, give each tool a distinct name and description, and consider tool search or deferred loading features where providers offer them.
Deeper: the Dubai logistics monitor (illustrative)
Let's deepen the Dubai logistics startup example with illustrative numbers. Forty real customer messages per month, in English and Arabic, run against two providers. For three months both scored within a couple of points on tool selection. Then, after a model update, one provider started calling reschedule delivery for messages that were really complaints, and accuracy dropped several points on Arabic messages. Because routing was configuration, they switched the Arabic route to the other provider the same afternoon and opened a ticket with the vendor. A month later, after a prompt tweak for the affected model, both were close again. No tool code changed at any point.
Watch me do it: order status loop
Watch me do it. Let's trace the Claude part of the order status example. The canonical schema has one property, order id, required, with no extra properties, and the tool definition adds strict true. The first call sends the question and the tool list. Claude replies with stop reason tool use. My while loop appends the assistant content, then builds a result for each tool use block: it calls get order status with the block's input and wraps the JSON output in a tool result with the matching id. I append all results as one user message and call again. This time the stop reason is end turn, so I print the text. The OpenAI loop is the same shape, except I extend my input list with the model's output items, parse the arguments string with json loads, and send function call output items. Gemini appends the model's content, then function response parts in a tool role message.
Recap + try this now
Quick recap. The loop is universal: define, call, execute, return, continue. Claude uses tool use and tool result blocks. OpenAI uses function call items with string arguments. Gemini uses function calls and function responses. Keep one canonical schema, validate in code, and cap every loop. Try this now: take one of your own tools, implement it against two providers with the canonical schema approach, add a step cap and an error path, and compare accuracy on ten realistic requests.
Key takeaways
- Tool calling follows the same loop everywhere: define, model calls, you execute, return results, continue.
- Claude uses tool_use/tool_result blocks; OpenAI uses function_call items with JSON-string arguments and function_call_output; Gemini uses function_calls and function responses.
- Keep one canonical JSON Schema per tool and generate provider formats from it.
- Return all parallel results together, send errors back as readable results, and cap loops.
- Validate arguments and permissions in code; hosted server-side tools vary by provider and platform.
Try it
Implement one of your own tools against two providers using the canonical-schema approach, add a step cap and an error path, and compare tool-selection accuracy on ten requests.