SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

OpenAI Makes Powerful 'GPT-4 Turbo with Vision' API Generally Available

—Making it easy to create apps that read and recognize images—

OpenAI has made the API for its powerful language model, 'GPT-4 Turbo with Vision,' generally available. This allows companies and developers to easily integrate advanced language processing and image recognition capabilities into their applications.

GPT-4 Turbo with Vision combines the image and audio upload capabilities of GPT-4, announced last September, with the faster GPT-4 Turbo model unveiled at OpenAI's developer conference in November.

This model offers many benefits, including significantly improved processing speeds, an input context window of up to 128,000 tokens (equivalent to approximately 300 pages of text), and affordable pricing for developers.

API requests can utilize the model's image recognition and analysis capabilities through JSON-formatted text and function calling. This allows developers to generate JSON code snippets that automate actions such as sending emails, making purchases, or posting online. However, OpenAI strongly recommends building user confirmation flows before executing actions that have real-world consequences.

Several startups are already leveraging GPT-4 Turbo with Vision.

・Healthify: A health and fitness app that provides nutritional analysis and recommendations when users upload photos of their meals.

・TLDraw: A UK-based startup that generates website applications based on web screen images and operational specifications drawn on a whiteboard.

For more details, please refer to the original article provided by OpenAI.

[Source]

https://platform.openai.com/docs/guides/vision

[Read Aloud]
VOICEVOX Shikoku Metan/No.7

いいなと思ったら応援しよう!