BMW is proposing a car that could summarize a friend’s message, imitate their voice and animate their contact photo to deliver it. The familiar face would be speaking an AI rewrite, complete with an expression based on how the software thinks the sender feels. That is a substantial step beyond having the dashboard read a text aloud.
The system appears in a patent filing discovered by Carmoses. BMW’s stated aim is to make incoming information quicker to absorb and potentially reduce driver distraction. Its most ambitious option also introduces a difficult distinction for anyone listening. The voice could belong to someone they know, while the wording and delivery come from the car’s software.
Nothing in the filing names a model but we used the iX3 for these renders because BMW has announced it as the first U.S. model for its Alexa+ powered Intelligent Personal Assistant.
A Familiar Face Reads A Rewrite
The process starts with information arriving through a connected device or the vehicle’s own communications system. That could be a text, email, voice recording, image or video. A generative AI model extracts the meaning and produces a shorter version for the speakers or a display. BMW does not specify a required button press or voice command to initiate playback.
The practical problem is recognizable. A notification showing only the beginning of a message may spend its limited space on a greeting and a name, leaving the useful part out of view. BMW proposes condensing the essential content instead, potentially into a single line.
Sender imitation is an optional extension. Given suitable voice information and a profile picture, the software could generate a video in which the pictured person appears to speak the summary. Their expression could change to match the emotion inferred from the original communication.

That output would be generated, rather than a recording of the person delivering those words. BMW mentions telling occupants that they are hearing a summary, but makes that disclosure optional. It supplies no accuracy results for the rewritten content or the emotional interpretation.
The Useful Part Is The Next Step
The proposed summary can separate what happened, how the sender seems to feel and what the recipient needs to do. BMW gives the example of an instruction to be home by 10 p.m. If the requested action includes a location, the system could pass it to navigation as a destination or an addition to the route.

That handoff has a clearer practical purpose than an animated portrait. It connects a request to a task the car can help complete. The emotional component is less straightforward. BMW suggests choosing an emoji, altering a photograph or changing the background color, with red representing danger and green representing okay. Those cues would present the software’s interpretation of the message, not independent knowledge of the sender’s state of mind.
The inputs could also include a phone conversation in the car or a discussion between passengers. BMW does not describe continuous recording. It allows the intended recipient’s identity to influence the summary or its presentation, using clues such as the destination device or a greeting addressed to a particular person.
Familiarity Needs Clear Boundaries
BMW first presented its emotionally expressive i Vision Dee concept in 2023, so the broader ambition to make digital interaction feel personal has a history. Earlier this year, it announced an Alexa+ based assistant capable of handling conversational requests and connecting answers to vehicle functions such as navigation.
Message reduction also predates this proposal. Hyundai and Kia previously described a system that adjusts spoken summaries to the driver’s state and can delay playback when driving demands greater attention. These are different approaches to the same problem of deciding how much information belongs in the cabin at a given moment.
BMW allows simpler processing to happen onboard and more complex work to run on an external server. It also proposes keeping passenger data out of model training and limiting its use to privately presenting the information. Those provisions leave unanswered questions about sender permission, data retention and when animated video would be available to a driver.
The filing offers no measured distraction benefit, and a patent does not guarantee production. The strongest case here is for a concise, accurate recap that preserves a requested action. Putting that recap into a familiar face and voice needs a higher standard. Clear disclosure that the delivery is generated should be mandatory before this becomes a feature drivers are expected to trust.


