
Lucia Ferrante
My beat is the seam between a model and the application around it, which is where most integration problems live. Structured output is the first problem. There are three mechanisms and they differ in how much freedom the model has. Free-form text asking for JSON gets you JSON most of the time. A tool or function definition with a declared schema constrains the arguments. A native structured output mode gives the strongest guarantee, and where it exists it is worth preferring. On every page I recommend validating the response against the schema anyway, because a guarantee at the provider is not a guarantee in your error budget. Refusals and truncated responses arrive as malformed output, and code that assumes a field exists will read undefined rather than handling the case. That is where retries help, and where blind retries hurt by doubling cost while hiding a systematic problem. Model choice is an architectural decision rather than a per-call one. Different models differ in latency, cost, context length and behaviour on the same prompt, and an abstraction layer that hides that choice also hides the ability to move. I write these pages so the trade-off stays visible. Streaming changes the failure surface. A stream can fail halfway, and an application that has already written to the user needs a way to recover that does not duplicate the partial output. Token-by-token output also changes what validation can do, since the structure is incomplete until the last fragment arrives. I close on observability, because an integration you cannot see is an integration you cannot debug: log the prompt, the model version and the response, with the identifiers you need to trace one call across services.
About ToolSura
ToolSura offers 80+ free, privacy-first online tools that run 100% in your browser — no uploads, no logins. Learn more about our mission →