Forget Artificial General Intelligence (AGI) – the big impact is already here and it’s called AI agents
...
The Rabbit r1 was a big hit at the CES show this year. It’s affordably priced, it allowed direct access to an AI engine. But the focus on the device may have missed the real bombshell that Rabbit r1 dropped.
In the video, CEO Jesse Lyu teaches the AI to book a vacation, but he’s not visiting the web-sites, nor is he clicking on any web pages. The “controller” that oversees the interaction of agents and websites is doing all the work. More than that, it’s learning from this interaction how to do this series of tasks to achieve a particular result. After its first training, he is able to just say, “plan a trip” with a few comments and it executest the planning and comes back with appropriate options.
Think about that. There is no need for interfaces or APIs. There is no need for the user to have standalone apps. Once trained on a similar task, the AI system (not the device) is able to learn to do everything necessary – including operating a browser or even a computer – on its own, without intervention, to achieve a specific outcome. Once it’s learned that task, it can adapt it and even modify it. ...
You might remember the scandal that happened when a company called Cambridge Analytics collected Facebook data and claimed they could use likes and dislikes to predict everything from your voting preferences to your sexual preferences. That’s going to seem “so 2016” when you consider that we might all have an agent that no longer has to predict – it will know what you are going to do.
That creates a huge dilemma in terms of privacy. Where will that be stored? Who will “own” it and control it? ...
For those not familiar with it, Metcalfe’s law, simply stated, says that the value of a network is the square of its nodes. Or to put it simply, once you get a critical mass of users in any platform, it becomes really hard for anyone else to compete. ...
How strong is Metcalfe’s law. It is so strong that even Elon Musk has not been able to totally kill X/Twitter. Millions have left for other platforms, but even so, few have actually deleted their Twitter account and totally moved on. Despite many new startups trying to supplant Twitter, as of yet, no-one has. The value of their network is still dwarfed by Twitter.
X/Twitter may lose enough money to destroy itself, but it hasn’t lost enough people. That’s the power of Metcalfe’s law. ...
See the full story here: https://www.itworldcanada.com/article/forget-artificial-general-intelligence-agi-the-big-impact-is-already-here-and-its-called-ai-agents/558707
Meta’s new AI model learns by watching videos
Meta’s AI researchers have released a new model that’s trained in a similar way to today’s large language models, but instead of learning from written words, it learns from video.
LLMs are normally trained on thousands of sentences or phrases where some of the words are masked, forcing the model to find the best words to fill in the blanks. In doing so they pick up a rudimentary sense of the world. Yann LeCun, who leads Meta’s FAIR (foundational AI research) group, has proposed that if AI models could use the same masking technique, but on video footage, they could learn more quickly. ...
The embodiment of LeCun’s theory is a research model called Video Joint Embedding Predictive Architecture (V-JEPA). It learns by processing unlabeled video and figuring out what probably happened in a certain part of the screen during the few seconds it was blacked out. ...
Meta’s next step after V-JEPA is to add audio to the video, which would give the model a whole new dimension of data to learn from—just like a child watching a muted TV then turning the sound up. ...
See the full story here: https://www.fastcompany.com/91029951/meta-v-jepa-yann-lecun
HOW GENERATIVE AI COULD ENABLE A NEW ERA OF FILMMAKING
1. Video generation: Generative video tools, such as Runway’s Gen-2and Pika, that use video diffusion models are capable of synthesizing novel video, creating short, soundless animations from text prompts, images or video.
2. Neural radiance fields (NeRFs): .... To create a NeRF, a neural network is trained on a simple recorded video from any camera or just a partial set of 2D images, meaning they don’t need to show all perspectives, sides or angles of the object or scene. The network is then able to generate a high-fidelity 3D representation of an object or scene by inferring unseen viewpoints, even those not captured in training data provided to the model. ...
3. Video avatars: Generative AI tools developed by Synthesia, Soul Machines and HeyGen can create entirely synthetic, photorealistic avatars that combine deepfake video and synthetic speech to precisely replicate a specific person’s appearance, voice, expressions and mannerisms. These unique personal AI avatars have been variously referred to as digital humans, twins, doubles or clones. ...
See the full story here: https://variety.com/vip/how-generative-ai-could-enable-a-new-era-filmmaking-1235898355/
Largest text-to-speech AI model yet shows ’emergent abilities’
Researchers at Amazon have trained the largest ever text-to-speech model yet, which they claim exhibits “emergent” qualities improving its ability to speak even complex sentences naturally. The breakthrough could be what the technology needs to escape the uncanny valley. ...
Here are examples of tricky text mentioned in the paper:
- Compound nouns: The Beckhams decided to rent a charming stone-built quaint countryside holiday cottage.
- Emotions: “Oh my gosh! Are we really going to the Maldives? That’s unbelievable!” Jennie squealed, bouncing on her toes with uncontained glee.
- Foreign words: “Mr. Henry, renowned for his mise en place, orchestrated a seven-course meal, each dish a pièce de résistance.
- Paralinguistics (i.e. readable non-words): “Shh, Lucy, shhh, we mustn’t wake your baby brother,” Tom whispered, as they tiptoed past the nursery.
- Punctuations: She received an odd text from her brother: ’Emergency @ home; call ASAP! Mom & Dad are worried…#familymatters.’
- Questions: But the Brexit question remains: After all the trials and tribulations, will the ministers find the answers in time?
- Syntactic complexities: The movie that De Moya who was recently awarded the lifetime achievement award starred in 2022 was a box-office hit, despite the mixed reviews.
PhilNote: in text-to-voice it also produces natural inflections.
See the full story here: https://techcrunch.com/2024/02/14/largest-text-to-speech-ai-model-yet-shows-emergent-abilities/?utm_source=substack&utm_medium=email
OpenAI introduces SORA
Sora can create a video up to 1 minute in length. It is currently being tested by Red Teams. The link has more info and multiple videos.
https://openai.com/sora?fbclid=IwAR2RELh3nFGj3iFaFYgZoHhI0AqbgwS0LSxK5vXx8bsjkTBBebRertDFa-4#capabilities
OpenAI CEO says ‘very subtle’ misalignments could make AI wreak havoc
OpenAI CEO Sam Altman on Tuesday warned the “very subtle societal misalignments” within artificial intelligence (AI) could cause things to “go horribly wrong,” while still expressing optimism about the technology’s benefits. ...
“I’m much more interested in the very subtle societal misalignments where we just have these systems out in society and through no particular ill intention, things just go horribly wrong,” added Altman, whose company created ChatGPT.
Altman, however, said he believes “things will go tremendously right,” if efforts are made to mitigate the downsides of AI technology. ...
See the full story here: https://thehill.com/policy/technology/4464982-openai-ceo-subtle-misalignments-ai-wreak-havoc/
MoviePass – Profitability Milestone Underscores the Significant Impact that Artificial Intelligence and Machine Learning Enhancements Have had on Managing Its Cinematic MarketplaceMoviePass –
... The MoviePass Cinematic Marketplace is an aggregator for the industry that uses AI and machine learning engines to improve attendance and performance. The proprietary credit system helps drive exploration of titles looking to compete against movies with much larger budgets. Based on internal member testing, MoviePass found that on average, there is a 40 percent shift to theater location offering the same movie for fewer credits. Members increase midweek attendance by 50 percent and go to an average of 2.4 different theater locations while using their MoviePass subscription. ...
MoviePass has the largest theater footprint of any subscription service featuring over 3500 locations across America and covering all 50 states with a reach of over 97 percent of the market. ...
See the full PR here: https://www.globenewswire.com/news-release/2024/02/13/2828293/0/en/MoviePass-Surpasses-One-Million-Movies-Seen-On-Its-New-Platform-and-Achieves-First-Profitable-Year-in-Company-History.html
Google and Anthropic Are Selling Generative AI to Businesses, Even as They Address Its Shortcomings
‘We’re not in a situation where you can just trust the model output,’ said Eli Collins, vice president of product management at Google DeepMind
... Jared Kaplan, co-founder and chief science officer of Anthropic, said the AI startup is working on a number of techniques that will reduce hallucinations, including building data sets where the model should respond to questions with, “I don’t know.” The idea is that the AI system can be trained to respond only when it has sufficient information, or will provide citations for its answers. ...
One solution is to make it easy for users to identify the sources of information that AI systems like its Gemini chatbot send back, said Eli Collins, vice president of product management at Google DeepMind. ...
Large language models, once they have been trained on certain data, can’t “delete” that information from what they have already learned, Kaplan said. ...
See the full story here: https://www.wsj.com/articles/google-and-anthropic-are-selling-generative-ai-to-businesses-even-as-they-address-its-shortcomings-ff90d83d
A.I. Can Help “Personalize” Policies to Reach the Right People
Combining the power of experiments with the potential of machine learning has tremendous implications for designing more effective public policy.
... Ultimately, the researchers concluded that the most effective policy is to target nudges to the middle of the group — students who are neither the most nor least likely to reapply for financial aid. At either end of this spectrum, the power of the nudge weakens, particularly for those who are the least likely to apply for aid. ...
The machine learning model, in essence, can help mitigate concerns of external validity by fitting the results from one population to another distinct population.
This hybrid approach has the potential to make experimentation less expensive by supporting faster iteration. As an experiment is running, machine learning can discern what works and suggest ways to fine-tune interventions in real time for maximum impact.
For policymakers, this adaptable, targeted process provides the ability to move beyond catchall approaches that are often costly and marginally effective. ...
The method, Spiess says, also illuminates blind spots — such as the people who are left behind by certain interventions.
See the full story here; https://www.gsb.stanford.edu/insights/ai-can-help-personalize-policies-reach-right-people
Disney’s Newest Robot Demonstrates Collaborative Cuteness
How two robots can be much more capable than one
... Having robots work together–especially if they have complementary skill sets–can open up some exciting opportunities, especially in the entertainment robotics space. At Walt Disney Imagineering, our research and development teams have been working on this idea of collaboration between robots, and we were able to show off the result of one such collaboration in Shanghai this week, when a little furry character interrupted the opening moments for the first-ever Zootopia land.
Our newest robotic character, Duke Weaselton, rolled onstage at the Shanghai Disney Resort for the first time last December, pushing a purple kiosk and blasting pop music. As seen in the video below, the audience got a kick out of watching him hop up on top of the kiosk and try to negotiate with the Chairman of Disney Experiences, Josh D’Amaro, for a new job. ...
Watch the video here: https://www.youtube.com/watch?v=ZRJFYlldkBw&t=361s (PhilNote: the robot is the better actor)
What might not be obvious at first is that the moment you just saw was enabled not by one robot, but by two. Duke Weaselton is the star of the show, but his dynamic motion wouldn’t be possible without the kiosk, which is its own independent, actuated robot. ...
The character is an expressive, bipedal robot with an exaggerated, animated motion style. It looks fantastic, but it’s not optimized for robust, reliable locomotion. The kiosk, meanwhile, is a simple wheeled system that behaves in a highly predictable way. ...
Because the character robot is tough and because there is some flexibility built into its motors and joints, small errors in placement and pose don’t create big problems like they might for a more conventional robot. The character can lean on the motorized kiosk to create the illusion that it is pushing it across the stage. The kiosk then uses a winch to hoist the character onto a platform, where electromagnets help stabilize its feet. ...
Watch the robots train here: https://www.youtube.com/watch?v=VFNruidzlZw&t=54s
See the full story here: https://spectrum.ieee.org/disney-robot-2666681104
Pages
- About Philip Lelyveld
- Mark and Addie Lelyveld Biographies
- Presentations and articles
- Trustworthy AI – A Market-Driven approach
- Tufts Alumni Bio