OpenAI has rolled out significant updates to its ChatGPT application, enhancing its versatility and interactivity. These updates introduce two exciting features for users:
ChatGPT now supports voice interaction, featuring five lifelike synthetic voices. Users can engage with ChatGPT by speaking to it, and it responds in real-time. ChatGPT can answer questions related to images. Users can upload pictures and inquire about descriptions or information.
These enhancements come on the heels of the recent announcement that ChatGPT will be integrated with DALL-E 3, OpenAI’s image generation model, enabling the chatbot to create images.
The voice interaction feature operates using two models: Whisper converts spoken words into text, while a new text-to-speech model transforms ChatGPT’s responses into spoken words. OpenAI trained these synthetic voices based on real actors’ voices to achieve a more natural and lifelike sound. Future plans may include allowing users to create custom voices.
OpenAI is also sharing its text-to-speech model with other companies, such as Spotify, which uses it to translate celebrity podcasts into multiple languages.
These updates highlight OpenAI’s ability to swiftly transform experimental models into practical and user-friendly products. ChatGPT Plus, the premium version of the app, now combines GPT-4 and DALL-E, positioning it as a competitor to voice assistants like Siri, Google Assistant, and Alexa.
The image recognition feature enables users to upload images and seek information from ChatGPT, a feature already utilized by applications like Be My Eyes, which caters to individuals with visual impairments.
OpenAI remains cautious about potential risks and remains committed to addressing misuse while prioritizing user safety. These updates enhance ChatGPT’s utility and user-friendliness, offering a richer and more interactive experience for users.


