OpenAI Makes Powerful 'GPT-4 Turbo with Vision' API Generally Available
—Making it easy to create apps that read and recognize images—
OpenAI has made the API for its powerful language model, 'GPT-4 Turbo with Vision,' generally available. This allows companies and developers to easily integrate advanced language processing and image recognition capabilities into their applications.
GPT-4 Turbo with Vision combines the image and audio upload capabilities of GPT-4, announced last September, with the faster GPT-4 Turbo model unveiled at OpenAI's developer conference in November.
This model offers many benefits, including significantly improved processing speeds, an input context window of up to 128,000 tokens (equivalent to approximately 300 pages of text), and affordable pricing for developers.
API requests can utilize the model's image recognition and analysis capabilities through JSON-formatted text and function calling. This allows developers to generate JSON code snippets that automate actions such as sending emails, making purchases, or posting online. However, OpenAI strongly recommends building user confirmation flows before executing actions that have real-world consequences.
Several startups are already leveraging GPT-4 Turbo with Vision.
・Healthify: A health and fitness app that provides nutritional analysis and recommendations when users upload photos of their meals.
The @healthifyme team built Snap using GPT-4 Turbo with Vision to give users nutrition insights through photo recognition of foods from around the world. pic.twitter.com/jWFLuBgEoA
— OpenAI Developers (@OpenAIDevs) April 9, 2024
・TLDraw: A UK-based startup that generates website applications based on web screen images and operational specifications drawn on a whiteboard.
Make Real, built by @tldraw, lets users draw UI on a whiteboard and uses GPT-4 Turbo with Vision to generate a working website powered by real code. pic.twitter.com/RYlbmfeNRZ
— OpenAI Developers (@OpenAIDevs) April 9, 2024
For more details, please refer to the original article provided by OpenAI.
[Source]
https://platform.openai.com/docs/guides/vision
[Read Aloud]
VOICEVOX Shikoku Metan/No.7
