Meta SAM 3 Released: Video Tracking, 3D & Open Weights!
Quick Overview
Meta released Segment Anything Model 3 (SAM 3), a unified model for detection, segmentation, and tracking in images and videos using text, exemplar, and visual prompts, alongside the Segment Anything Playground for interactive experimentation, which includes features like video cutouts, image cutouts, 3D scene creation, and 3D body reconstruction.
Key Points: Meta announced Segment Anything Model 3 (SAM 3), a unified model for detection, segmentation, and tracking of objects in both images and videos. SAM 3 supports text, exemplar, and visual prompts, building on the capabilities of Segment Anything 2. The release includes the Segment Anything Playground, an interactive platform where users can experiment with SAM 3 features like creating video cutouts, image cutouts, 3D scenes, and 3D bodies. SAM 3 capabilities will soon enable new effects in Instagram's video creation app, Edits, and will be available on Vibes on the Meta AI app and meta.ai. SAM 3D, a suite for 3D objects and human reconstruction from a single image, is also being shared, setting a new standard for grounded 3D reconstruction. The model demonstrates state-of-the-art performance across various benchmarks, including Concept Segmentation, Visual Segmentation, Counting, and Reasoning Segmentation. The video showcases SAM 3's ability to segment and track objects (like people, elephants, zebras, and jets) in complex scenes and video sequences using text prompts.
Context: This video announces the release of Meta's next-generation visual foundation model, Segment Anything Model 3 (SAM 3), which extends the SAM series to handle object tracking in videos and introduces advanced prompting capabilities like text and visual prompts. The announcement also covers the introduction of the Segment Anything Playground, a web platform allowing users to test these new capabilities immediately across various tasks such as video cutouts and 3D reconstruction.
Detailed Analysis
Meta officially announced Segment Anything Model 3 (SAM 3), an evolution of the SAM series that unifies detection, segmentation, and tracking across images and videos using text, exemplar, and visual prompts. This new iteration is noted to be 50% faster for positive prompts in challenging, fine-grained domains compared to previous models. Alongside SAM 3, Meta released the Segment Anything Playground, an interactive environment where users can test features like creating video cutouts (03:51), image cutouts (05:19), creating 3D scenes (06:21), and generating 3D bodies (06:56) from 2D inputs. The announcement highlights practical applications, such as integrating SAM 3 into Instagram's Edits app for new video effects and powering Facebook Marketplace's 'View in Room' feature via SAM 3D. The video provides extensive demonstrations, showing SAM 3 successfully tracking objects like people in a park (00:01), fish underwater (00:09), penguins (00:14), dogs (00:19), and jets in formation (09:40), all driven by text prompts like "people," "zebra," or "jet." Furthermore, the video details the SAM 3 data engine workflow, which relies on a hybrid human and AI verification loop for quality control, and showcases SAM 3D's ability to generate 3D models (06:44) from single images. Benchmarks confirm SAM 3 achieves state-of-the-art results across numerous segmentation tasks (02:58).