AI at Meta’s cover photo
AI at Meta

AI at Meta

Research Services

Menlo Park, California 1,119,356 followers

Together with the AI community, we’re pushing boundaries through open science to create a more connected world.

About us

Through open science and collaboration with the AI community, we are pushing the boundaries of artificial intelligence to create a more connected world. We can’t advance the progress of AI alone, so we actively engage with the AI research and academic communities. Our goal is to advance AI in Infrastructure, Natural Language Processing, Generative AI, Vision, Human-Computer Interaction and many other areas of AI enable the community to build safe and responsible solutions to address some of the world’s greatest challenges.

Website
https://ai.meta.com/
Industry
Research Services
Company size
10,001+ employees
Headquarters
Menlo Park, California
Specialties
research, engineering, development, software development, artificial intelligence, machine learning, machine intelligence, deep learning, computer vision, engineering, computer vision, speech recognition, and natural language processing

Updates

  • AI at Meta reposted this

    Muse Image is now available on Meta Model API and priced for production volumes at $0.01/image.  Three reasons it's worth trying out: - It's agentic. Muse Image reasons before it renders. It can plan a complex request, search the web for real references, and can write and run code for precise elements like charts and QR codes. - Quality holds across edits. Generate from a description or edit what you already have without quality drift. Text-to-image, single-image editing and multi-image editing all live in one model, so there's no multi-step pipeline to stitch together. - The economics work at scale. Production-grade image generation at $0.01/image. Workloads that were often prohibitive at frontier pricing become routine. Try it now on whatever you’re building. Read more at https://bit.ly/4gzEIWt

  • View organization page for AI at Meta

    1,119,356 followers

    Muse Spark 1.2 supports a broad range of multimodal tasks, from turning visuals into working code to translating perception into physical action. It also brings robust audio-visual understanding to enable video-heavy workflows common in real-world enterprise use. Today, we’re sharing new evals and demos that illustrate the breadth of the model’s visual understanding and reasoning capabilities. A few highlights: 1️⃣ Tool use: The model’s multimodal gains are most pronounced with tool use. The model inspects visual inputs more closely and incorporates what it finds into its reasoning. 2️⃣ Visual coding: It builds artifacts, like web pages and games, from images or video. It translates visual layout, hierarchy, and style into working code, evaluating correctness based on actual rendering and behavior while using a continuous self-improvement loop to refine its outputs. 3️⃣ Spatial Reasoning: The model combines deep video understanding and dense captioning with a suite of agentic tools to move beyond passive observation toward autonomous orchestration of complex workflows. For example, Muse Spark 1.2 can parse multimodal observations and calls tools to guide a robot to navigate in an unstructured environment (see a demo in the link below). 4️⃣ Audio-Visual Workflows: It combines deep video understanding and dense captioning with a suite of agentic tools (web development, real-time search, and spatial grounding) to translate visual input into actionable outputs. Read the full research deep dive to see more evals and demos: go.meta.me/multimodal

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • View organization page for AI at Meta

    1,119,356 followers

    Introducing Muse Glimmer, an open-weight 30-billion-parameter model optimized for always-on local agent workflows. Keeping with our long tradition of sharing fundamental AI research, we’re releasing the weights under a permissive Apache 2.0 license. Muse Glimmer is small enough to run on consumer hardware like a Mac or PCs with performant GPUs, supporting local agents and function calling, local coding, and LLM-as-a-judge evaluation. The model delivers strong performance on key agentic use cases and benchmarks compared with leading models in its size category. Most foundation model deployments still depend on cloud infrastructure and network access. Running locally means AI that works anywhere, with or without a connection. Muse Glimmer is optimized for exactly that. 🔗 Read the research blog: https://lnkd.in/enw7a4St 🔗 Download Muse Glimmer on Hugging Face: https://lnkd.in/eK6x9MQc 🔗 Find resources: https://lnkd.in/eQ9GWHxE

    • No alternative text description for this image
    • No alternative text description for this image
  • View organization page for AI at Meta

    1,119,356 followers

    To understand whether we're making genuine progress on reasoning, we entered our AI models in five STEM Olympiad competitions. The results: 🏅 Asian Physics Olympiad (APhO): Perfect score, theory exam 🏅 International Physics Olympiad (IPhO): Perfect score, theory exam 🥇 International Mathematical Olympiad (IMO): Gold medal 🥇 International Chemistry Olympiad (IChO): Gold-medal-level performance 🥇 Romanian Masters of Mathematics (RMM): Gold-medal-level performance These problems were exceptionally difficult, and demanded deep chains of reasoning, creative insight, and flawless formal argumentation. To test pure reasoning capability, we disallowed all tool use, meaning no search, no coding, and no calculator. We have deep admiration for the contestants and committees behind these competitions, and are grateful for their support in enabling our participation.

    • No alternative text description for this image
  • View organization page for AI at Meta

    1,119,356 followers

    Introducing Muse Code (beta), a terminal coding agent built for long-horizon software engineering, powered by our new Muse Spark 1.2 model. Muse Code plans, implements, and verifies complex changes across large repositories. It coordinates persistent sub-agents to solve difficult problems faster, more accurately, and with less intervention. Key Technical Highlights: 1️⃣ Async Background Agents: Muse Code operates with a simple agent loop enhanced by persistent background agents that stay active throughout each session rather than spawning per task, reducing latency, avoiding redundant information gathering, and minimizing the need for steering on difficult, multi-step tasks. 2️⃣ Runtime Design: Every model call, tool invocation, edit, and approval is logged to a single source of truth. This guarantees replay-exact state reconstruction, allowing long-running tasks to resume seamlessly after system failures. 3️⃣ Co-Trained Model: Muse Spark 1.2 was co-trained alongside Muse Code using rejection-sampled harness trajectories and context compaction. A self-improvement loop using Muse Spark 1.1 generated environments and auto-graded solutions to improve instruction-following precision. 4️⃣ Long-Horizon Impact: Iterative execution over 1,000+ tool calls (up to 24 hours) to autonomously optimize KDA and MLA GPU kernels on NVIDIA Hopper GPUs, achieving substantial improvements over baseline implementations. Muse Spark 1.2 is available today in Muse Code and in Meta Model API, with expanded global access and many new features on the horizon. Learn more: https://lnkd.in/eH59rFCQ

    • No alternative text description for this image
    • No alternative text description for this image
  • View organization page for AI at Meta

    1,119,356 followers

    The U.S. Department of Energy (DOE)’s Genesis Mission aims to accelerate scientific discovery with AI across DOE's national laboratories. One of its flagship projects, SYNAPS-I, is putting that vision into practice with Meta's open-source SAM 3 and DINOv3 models. Researchers at Berkeley Lab, which is leading SYNAPS-I in partnership with other leading labs, fine-tuned SAM 3 and DINOv3 on scientific imaging data and deployed them across 300 A100 GPUs. DINOv3 identifies structures within images while SAM 3 draws precise pixel-level boundaries. Together, they deliver a fully labeled 3D volume back to scientists in ~15 minutes, a process that previously took a month of manual work. Learn more about their work here: https://go.meta.me/633fe4

    • No alternative text description for this image
  • View organization page for AI at Meta

    1,119,356 followers

    We’re excited to introduce Muse Spark 1.1, a significant upgrade to the first Muse Spark model we released earlier this year. Muse Spark 1.1 is a multimodal reasoning model built for agentic tasks, with major gains in coding, tool and computer use, and multimodal understanding. The model is available now in "Thinking" mode in the Meta AI app and on meta.ai. Along with this release, we are launching a public preview of the new Meta Model API where developers can access and build with Muse Spark 1.1. Learn more: https://go.meta.me/646233

    • No alternative text description for this image
  • View organization page for AI at Meta

    1,119,356 followers

    Muse Image is our most advanced image generation model yet, deploying an agentic workflow that co-plans with Muse Spark to search the web, call tools, and reflect on its own outputs to refine them before delivering the final image. A few things that Muse Image can do: 1️⃣ Write and execute code to nail precise details like plots and QR codes, and team up with Muse Spark to produce websites with embedded images and playable visual games. 2️⃣ Search the web to ground generated images in factual and real-time information and visual references. 3️⃣ Compose elements from many input reference images in the prompt, including people, objects, clothing, styles, and environments. It supports interleaving text and images inline in prompts for complex image compositions. You can try Muse Image today in the Meta AI app and web, as well as in Instagram Stories and WhatsApp – starting in limited countries with more locations on the way. Learn more about Muse Image: https://go.meta.me/080c53

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
      +1
  • View organization page for AI at Meta

    1,119,356 followers

    Introducing Muse Image and Muse Video, the first media generation models developed by Meta Superintelligence Labs. Muse Image is our most advanced image generation model yet. It follows instructions faithfully, edits with precision, composes from multiple references, and draws on Instagram for social context. It also brings agentic tool use capabilities to image generation and integrates with Muse Spark. You can try Muse Image in the Meta AI app and web, as well as in Instagram Stories and WhatsApp – starting in limited countries with more locations on the way. Today we’re also previewing Muse Video, which is built upon the same pretraining base as Muse Image. It offers competitive performance in prompt adherence, visual fidelity, and temporal consistency. We’re investing in areas with current performance gaps, such as audio-video synchronization and physically accurate fast motion. Learn more about both models: https://go.meta.me/080c53

Affiliated pages

Similar pages