Welcome to Science Sessions, the PNAS podcast program. Listen to brief conversations with cutting-edge researchers, Academy members, and policymakers as they discuss topics relevant to today's scientific community. Learn the behind-the-scenes story of work published in PNAS, plus a broad range of scientific news about discoveries that affect the world around us.
…
continue reading
A daily update on the latest AI Research Papers. We provide a high level overview of a handful of papers each day and will link all papers in the description for further reading. This podcast is created entirely with AI by PocketPod. Head over to https://pocketpod.app to learn more.
…
continue reading
Talking to people who use maths in their work. Aiming to encourage further uptake of maths at A-level and beyond. Hosted by Peter Rowlett and Katie Steckles.
…
continue reading
This is a feed of pages for Mike Godfrey
…
continue reading
1
Truth from New Thought: Real Personal Transformation - Discover Pathways, If You Want to Change Your Life - Insights and Strategies for Growth
Ekarach Chandon
The Series Book to Help You Understand 'Humanity' and 'Self' on a Deeper Level... Humans have always been curious, exploring the universe and its mysteries. But, an essential question arises: What should be the first subject of our study? It seems imperative that understanding humanity should be our priority. After all, knowing others and ourselves is foundational before venturing into the unknown realms of the universe. This series of books delves into the essence of humanity: "Arn'-Gone'-C ...
…
continue reading
Cryptography FM is a regular podcast with news and a featured interview covering the latest developments in theoretical and applied cryptography. Whether it's a new innovative paper on lattice-based cryptography or a novel attack on a secure messaging protocol, we'll get the people behind it on Cryptography FM.
…
continue reading
The Journal of Proteome Research integrates the fields of chemistry, mathematics, applied physics, biology, and medicine in order to better understand the function of proteins in biological systems.
…
continue reading
Nodycast is a lively podcast discussing the theory, techniques and latest innovations in nonlinear dynamics, and its applications to systems of all kinds. This includes almost everything under the sun such as mechanical, structural, electrical, chemical, thermo-fluid, ecological, economic, epidemiological, biological and chemical systems. It is hosted by Dr. 'Nat' C. Nataraj, Moritz Professor at Villanova University and Senior Editor for Nonlinear Dynamics, a Springer-Nature journal.
…
continue reading
Mathematical Philosophy - the application of logical and mathematical methods in philosophy - is about to experience a tremendous boom in various areas of philosophy. At the new Munich Center for Mathematical Philosophy, which is funded mostly by the German Alexander von Humboldt Foundation, philosophical research will be carried out mathematically, that is, by means of methods that are very close to those used by the scientists. The purpose of doing philosophy in this way is not to reduce p ...
…
continue reading
The Department of Statistics at Oxford is a world leader in research including computational statistics and statistical methodology, applied probability, bioinformatics and mathematical genetics. In the 2014 Research Excellence Framework (REF), Oxford's Mathematical Sciences submission was ranked overall best in the UK. This is an exciting time for the Department. We have now moved into our new home on St Giles and we are currently settling in. The new building provides improved lecture and ...
…
continue reading
If you're searching for an authentic, career-focused podcast, designed by STEM and healthcare professionals for STEM and healthcare professionals, you're in the right place! The White Coat White Collar® Podcast is on a mission to demystify the career landscape so STEM and healthcare students, graduates, and professionals can find the path best suited to them. In each bi-weekly episode, podcast host Dr. Aurellia Whitmore dives deep into the diverse career options in the science, technology, e ...
…
continue reading
The Last Theory is an easy-to-follow exploration of what might be the last theory of physics. In 2020, Stephen Wolfram launched the Wolfram Physics Project to find the elusive fundamental theory that explains everything. On The Last Theory podcast, I investigate the implications of Wolfram's ideas and dig into the details of how his universe works. Join me for fresh insights into Wolfram Physics every other week.
…
continue reading
A podcast about virtual reality, hosted by Ctrl V co-hosts Josh Brooks & Ben Parent.
…
continue reading
Podcast by Psychology In Action Podcast
…
continue reading
1
Diffusion Forcing to Expert Tuning, Structured Planning, Vision-Language Models, and Tabular ML Benchmarks
11:34
11:34
Play later
Play later
Lists
Like
Liked
11:34
Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionLet the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language ModelsPlanetarium: A Rigorous Benchmark for Translating Text to Structured Planning LanguagesInternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Co…
…
continue reading
1
Advancing AI's Mathematical Reasoning: WE-MATH, ROS-LLM Framework, Autoregressive Image Generation
10:36
10:36
Play later
Play later
Lists
Like
Liked
10:36
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?ROS-LLM: A ROS framework for embodied AI with task feedback and structured reasoningMMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient EvaluationLiteSearch: Efficacious Tree Search for LLMWavelets Are All You Need for Autoregressive Image…
…
continue reading
1
Success Ratio: The Mathematics of Success Hidden in the Universe
2:58
2:58
Play later
Play later
Lists
Like
Liked
2:58
Begin with 'Success Ratio' to enter the realm of dispelling ignorance and achieving true self-transformation according to the natural laws of the universe. Now available for purchase on Amazon.com "Asking even a slightly wrong question about life can lead to completely different paths." The Search Ends Here DON'T LET YOUR SUBCONSCIOUS DISMISS THIS …
…
continue reading
1
Persona-Driven Data Synthesis, Enhancing Medical MLLMs, Robot Learning, Knowledge Distillation in LLMs, Text to 3D Gaussian Revolution
11:24
11:24
Play later
Play later
Lists
Like
Liked
11:24
Scaling Synthetic Data Creation with 1,000,000,000 PersonasHuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at ScaleLLaRA: Supercharging Robot Learning Data for Vision-Language PolicyDirect Preference Knowledge Distillation for Large Language ModelsGaussianDreamerPro: Text to Manipulable 3D Gaussians with Highly Enh…
…
continue reading
1
OMG-LLaVA: Unifying Vision and Language Understanding, Step-DPO for LLMs Mathematical Reasoning, MUMU's Multimodal Image Generation
12:15
12:15
Play later
Play later
Lists
Like
Liked
12:15
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and UnderstandingStep-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMsMUMU: Bootstrapping Multimodal Image Generation from Text-to-Image DataSimulating Classroom Education with LLM-Empowered AgentsSeaKR: Self-aware Knowledge Retrieval for Adaptive Retrieval …
…
continue reading
Inequitable wildfire smoke exposure in California Science Sessions are brief conversations with cutting-edge researchers, National Academy members, and policymakers as they discuss topics relevant to today's scientific community. Learn the behind-the-scenes story of work published in the Proceedings of the National Academy of Sciences (PNAS), plus …
…
continue reading
1
FineWeb Datasets, YouDream's 3D Animals, PDE-Solving Breakthrough, Noise-Conditioned Perception Alignment, Language Models' Continual Learning
11:02
11:02
Play later
Play later
Lists
Like
Liked
11:02
The FineWeb Datasets: Decanting the Web for the Finest Text Data at ScaleYouDream: Generating Anatomically Controllable Consistent Text-to-3D AnimalsDiffusionPDE: Generative PDE-Solving Under Partial ObservationAligning Diffusion Models with Noise-Conditioned PerceptionUnlocking Continual Learning Abilities in Language Models…
…
continue reading
1
BigCodeBench Challenges, Cambrian-1 Leap, D-MERIT's Evaluation, Long Context Breakthrough in Vision
11:06
11:06
Play later
Play later
Lists
Like
Liked
11:06
DreamBench++: A Human-Aligned Benchmark for Personalized Image GenerationBigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex InstructionsCambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsEvaluating D-MERIT of Partial-annotation on Information RetrievalLong Context Transfer from Language to Vision…
…
continue reading
1
LongRAG Breakthrough, LLMs as Judges, Transformer Memory Insights, Video Library AI, Democratizing Art Styles
10:14
10:14
Play later
Play later
Lists
Like
Liked
10:14
LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMsJudging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-JudgesComplexity of Symbolic Representation in Working Memory of Transformer Correlates with the Complexity of a TaskTowards Retrieval Augmented Generation over Large Video LibrariesStylebreeder: Exploring …
…
continue reading
1
Scaling In-Context Reinforcement Learning, ChartMimic's AI Benchmark, Multimodal Document Comprehension, Long Context Reasoning Challenges
10:36
10:36
Play later
Play later
Lists
Like
Liked
10:36
XLand-100B: A Large-Scale Multi-Task Dataset for In-Context Reinforcement LearningMake It Count: Text-to-Image Generation with an Accurate Number of ObjectsChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code GenerationNeedle In A Multimodal HaystackBABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Hay…
…
continue reading
1
Revolutionizing Vision and Language Models: Depth Prediction Breakthroughs, Pixel-Level Transformers, and Robotic Skill Learning
13:20
13:20
Play later
Play later
Lists
Like
Liked
13:20
Depth Anything V2An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual PixelsTransformers meet Neural Algorithmic ReasonersSamba: Simple Hybrid State Space Models for Efficient Unlimited Context Language ModelingOpenVLA: An Open-Source Vision-Language-Action ModelAlleviating Distortion in Image Generation via Multi-Resolut…
…
continue reading
Biodiversity and gentrification Science Sessions are brief conversations with cutting-edge researchers, National Academy members, and policymakers as they discuss topics relevant to today's scientific community. Learn the behind-the-scenes story of work published in the Proceedings of the National Academy of Sciences (PNAS), plus a broad range of s…
…
continue reading
1
NaRCan Revolutionizes Video Editing, Training-Free Video Generation, Recaptioning Web Images with LLaMA-3, Novel Data Synthesis Approach, Smartphone LLM Inference
11:33
11:33
Play later
Play later
Lists
Like
Liked
11:33
NaRCan: Natural Refined Canonical Image with Integration of Diffusion Prior for Video EditingMotionClone: Training-Free Motion Cloning for Controllable Video GenerationWhat If We Recaption Billions of Web Images with LLaMA-3?Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with NothingPowerInfer-2: Fast Large Language Model I…
…
continue reading
1
Revolutionizing Image Synthesis with TiTok, Multilingual Code Benchmark, Exploring GenAI Prompting Techniques,
10:53
10:53
Play later
Play later
Lists
Like
Liked
10:53
An Image is Worth 32 Tokens for Reconstruction and GenerationMcEval: Massively Multilingual Code EvaluationZero-shot Image Editing with Reference ImitationThe Prompt Report: A Systematic Survey of Prompting TechniquesTextGrad: Automatic "Differentiation" via Text
…
continue reading
1
LlamaGen's Image Revolution, Husky: The Multi-Step Reasoner, Vript's Video Breakthrough, VALL-E 2 Achieves Human Parity
10:46
10:46
Play later
Play later
Lists
Like
Liked
10:46
Autoregressive Model Beats Diffusion: Llama for Scalable Image GenerationHusky: A Unified, Open-Source Language Agent for Multi-Step ReasoningVript: A Video Is Worth Thousands of WordsLighting Every Darkness with 3DGS: Fast Training and Real-Time Rendering for HDR View SynthesisVALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text …
…
continue reading
1
Mixture-of-Agents, Benchmarking LLMs, and GenAI Arena Evaluation
11:06
11:06
Play later
Play later
Lists
Like
Liked
11:06
Mixture-of-Agents Enhances Large Language Model CapabilitiesWildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the WildCRAG -- Comprehensive RAG BenchmarkGenAI Arena: An Open Evaluation Platform for Generative ModelsLarge Language Model Confidence Estimation via Black-Box Access
…
continue reading
1
Enhancing AI Video and Image Generation, BitsFusion Quantization, Step-aware Optimization, Thought-Augmented Reasoning, and Single Forward Video Generation
11:39
11:39
Play later
Play later
Lists
Like
Liked
11:39
ShareGPT4Video: Improving Video Understanding and Generation with Better CaptionsBitsFusion: 1.99 bits Weight Quantization of Diffusion ModelStep-aware Preference Optimization: Aligning Preference with Denoising Performance at Each StepBuffer of Thoughts: Thought-Augmented Reasoning with Large Language ModelsSF-V: Single Forward Video Generation Mo…
…
continue reading
1
AI Papers Podcast Special Edition: Apple Intelligence & Ferret-UI
1:52
1:52
Play later
Play later
Lists
Like
Liked
1:52
Apple announced new Siri features and Apple Intelligence today, Interestingly, Apple already released a paper, titled "Ferret-UI," on how it all works - a multimodal vision-language model capable of understanding widgets, icons, and text on an iOS mobile screen, and reasoning about their spatial relationships and functional meanings. https://arxiv.…
…
continue reading
1
Block Transformers: Faster Inference, Mobile Device AI Agents, 3D-Image Generation, Low Latency TTS
10:41
10:41
Play later
Play later
Lists
Like
Liked
10:41
Block Transformer: Global-to-Local Language Modeling for Fast InferenceParrot: Multilingual Visual Instruction TuningMobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent CollaborationOuroboros3D: Image-to-3D Generation via 3D-aware Recursive DiffusionLiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autore…
…
continue reading
1
Seed-TTS, Decoding LLMs, Innovations in Text-to-Video, Self-Improving AI Preferences, and Refining Diffusion Models
11:10
11:10
Play later
Play later
Lists
Like
Liked
11:10
Seed-TTS: A Family of High-Quality Versatile Speech Generation ModelsTo Believe or Not to Believe Your LLMI4VGen: Image as Stepping Stone for Text-to-Video GenerationSelf-Improving Robust Preference OptimizationGuiding a Diffusion Model with a Bad Version of Itself
…
continue reading
1
MMLU-Pro: Next-Level Language Understanding, Tailored LLMs, High FPS Video Generation Innovation
11:30
11:30
Play later
Play later
Lists
Like
Liked
11:30
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding BenchmarkLearning Temporally Consistent Video Depth from Video Diffusion PriorsShow, Don't Tell: Aligning Language Models with Demonstrated FeedbackArtificial Generational Intelligence: Cultural Accumulation in Reinforcement LearningZeroSmooth: Training-free Diffuser Adaptati…
…
continue reading
1
Transformers and State-Space Models Unite, Multi-modal LLM Benchmark, Perplexity in Data Pruning, Advancing 4D Content Generation
10:23
10:23
Play later
Play later
Lists
Like
Liked
10:23
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityVideo-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video AnalysisPerplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference ModelsKaleido Diffusion: Improving Conditional Diffusion Models with Au…
…
continue reading
1
DITTO-2 Speeds Up Music AI, GECO's Quick 3D Generation, PLA4D's 4D Advances, DevEval's Real-World Code Benchmark, Parrot's LLM Application Efficiency
10:47
10:47
Play later
Play later
Lists
Like
Liked
10:47
AI Papers Podcast for 06/04/2024 DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music GenerationGECO: Generative Image-to-3D within a SECOndPLA4D: Pixel-Level Alignments for Text-to-4D Gaussian SplattingDevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code RepositoriesParrot: Efficient Serving of LLM-b…
…
continue reading
1
School enrollment during the COVID-19 pandemic
10:10
10:10
Play later
Play later
Lists
Like
Liked
10:10
School enrollment during COVID-19 Science Sessions are brief conversations with cutting-edge researchers, National Academy members, and policymakers as they discuss topics relevant to today's scientific community. Learn the behind-the-scenes story of work published in the Proceedings of the National Academy of Sciences (PNAS), plus a broad range of…
…
continue reading
1
Boosting Text Retrieval with CLIP Models, Rethinking Retrieval Augmented Generation, and Deciphering Human Behavior through MotionLLM
10:42
10:42
Play later
Play later
Lists
Like
Liked
10:42
AI Papers Podcast for 06/03/2024 Jina CLIP: Your CLIP Model Is Also Your Text RetrieverSimilarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered ThoughtsMotionLLM: Understanding Human Behaviors from Human Motions and VideosXwin-LM: Strong and Scalable Alignment Practice for LLMsMOFA-Video: Controllable Image Animati…
…
continue reading
1
Jonathan Gorard: the complete first interview
2:48:59
2:48:59
Play later
Play later
Lists
Like
Liked
2:48:59
I’ve heard from many of you that you’d like the whole of my conversation with Jonathan Gorard in a single podcast. So here it is, the complete first interview. These three hours are a brilliant exposition of Wolfram Physics from a figure whose contributions to the project are second to none. — Jonathan Gorard Jonathan Gorard at The Wolfram Physics …
…
continue reading
1
Bilingual LLM Transparency, T2V-Turbo's Video Generation, LLMs Surpassing Human Theory of Mind Performance, Advancements in LLM Attribution
8:47
8:47
Play later
Play later
Lists
Like
Liked
8:47
MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model SeriesT2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward FeedbackLLMs achieve adult human performance on higher-order theory of mind tasksNearest Neighbor Speculative Decoding for LLM Generation and AttributionZipper: A Multi-Tower Decoder Ar…
…
continue reading
1
Phased Consistency Model, 2-Stage Backpropagation, and the Future of 4D World Reconstruction
8:09
8:09
Play later
Play later
Lists
Like
Liked
8:09
Phased Consistency Model2BP: 2-Stage BackpropagationGFlow: Recovering 4D World from Monocular VideoInstruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction TuningLLaMA-NAS: Efficient Neural Architecture Search for Large Language Models
…
continue reading
1
Vision-Language Models, Arithmetic Transformers, Next-Gen Video Editing:
10:20
10:20
Play later
Play later
Lists
Like
Liked
10:20
An Introduction to Vision-Language ModelingTransformers Can Do Arithmetic with the Right EmbeddingsMatryoshka Multimodal ModelsI2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion ModelsZamba: A Compact 7B SSM Hybrid ModelLooking Backward: Streaming Video-to-Video Translation with Feature Banks…
…
continue reading
1
ConvLLaVA's Visual Compression, Efficient LLVM, Multilingual Aya 23, and AutoCoder's Code Mastery
11:11
11:11
Play later
Play later
Lists
Like
Liked
11:11
ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal ModelsMeteor: Mamba-based Traversal of Rationale for Large Language and Vision ModelsGrokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of GeneralizationAya 23: Open Weight Releases to Further Multilingual ProgressStacking Your Transformers: A Close…
…
continue reading
1
Revolution in Image Generation, Thermodynamic Gradient Descent, DMD2 for Fast Synthesis, Distributed Speculative Inference
10:56
10:56
Play later
Play later
Lists
Like
Liked
10:56
…
continue reading
1
Language Model Mysteries, Personalized Image Generation, Audio-Visual Transformer Innovations, DeepSeek-Prover, Dense Connector: MLLM Potential
10:31
10:31
Play later
Play later
Lists
Like
Liked
10:31
ReVideo: Remake a Video with Motion and Content ControlNot All Language Model Features Are LinearRectifID: Personalizing Rectified Flow with Anchored Classifier GuidanceVisual Echoes: A Simple Unified Transformer for Audio-Visual GenerationDeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic DataDense Connector for MLLMs…
…
continue reading
1
Transformer Linearity, Face-Adapter Diffusion Models, Cross-Layer Attention Shrinks LLMs, Image Generation Breakthrough
10:14
10:14
Play later
Play later
Lists
Like
Liked
10:14
Your Transformer is Secretly LinearDiffusion for World Modeling: Visual Details Matter in AtariFace Adapter for Pre-Trained Diffusion Models with Fine-Grained ID and Attribute ControlReducing Transformer Key-Value Cache Size with Cross-Layer AttentionOmniGlue: Generalizable Feature Matching with Foundation Model GuidancePersonalized Residuals for C…
…
continue reading
1
Infinite Video Generation, High-Rank Fine-Tuning, Modular LLMs with LoRA Libraries
9:18
9:18
Play later
Play later
Lists
Like
Liked
9:18
FIFO-Diffusion: Generating Infinite Videos from Text without TrainingMoRA: High-Rank Updating for Parameter-Efficient Fine-TuningOpenRLHF: An Easy-to-use, Scalable and High-performance RLHF FrameworkImp: Highly Capable Large Multimodal Models for Mobile DevicesOcto: An Open-Source Generalist Robot PolicyTowards Modular LLMs by Building and Reusing …
…
continue reading
1
Tailoring Language Models for Science, Scaling Laws in NLP, Grounded 3D-LLM Innovations, Efficient Large Model Inference
9:30
9:30
Play later
Play later
Lists
Like
Liked
9:30
INDUS: Effective and Efficient Language Models for Scientific ApplicationsObservational Scaling Laws and the Predictability of Language Model PerformanceGrounded 3D-LLM with Referent TokensLayer-Condensed KV Cache for Efficient Inference of Large Language ModelsDynamic data sampler for cross-language transfer learning in large language models…
…
continue reading
Emotional power of live music Science Sessions are brief conversations with cutting-edge researchers, National Academy members, and policymakers as they discuss topics relevant to today's scientific community. Learn the behind-the-scenes story of work published in the Proceedings of the National Academy of Sciences (PNAS), plus a broad range of sci…
…
continue reading
1
Chameleon's Multimodal Breakthrough, LoRA's Learning Efficiency, Many-Shot In-Context Learning, Object Detection Innovation, Text-to-3D Generation
10:29
10:29
Play later
Play later
Lists
Like
Liked
10:29
Chameleon: Mixed-Modal Early-Fusion Foundation ModelsLoRA Learns Less and Forgets LessMany-Shot In-Context Learning in Multimodal Foundation ModelsCAT3D: Create Anything in 3D with Multi-View Diffusion ModelsGrounding DINO 1.5: Advance the "Edge" of Open-Set Object DetectionDual3D: Efficient and Consistent Text-to-3D Generation with Dual-mode…
…
continue reading
1
Efficient Multimodality, Vision Suite's Custom Data, EEG Music Decoding Advances, Mobile Video Breakthrough
8:44
8:44
Play later
Play later
Lists
Like
Liked
8:44
ALPINE: Unveiling the Planning Capability of Autoregressive Learning in Language ModelsXmodel-VLM: A Simple Baseline for Multimodal Vision Language ModelBEHAVIOR Vision Suite: Customizable Dataset Generation via SimulationNaturalistic Music Decoding from EEG Data via Latent Diffusion ModelsNo Time to Waste: Squeeze Time into Channel for Mobile Vide…
…
continue reading
1
Transformer Models Beyond Scaling, Multilingual Image Synthesis, Advanced Text-to-Image Control
9:28
9:28
Play later
Play later
Lists
Like
Liked
9:28
VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion ModelsBeyond Scaling Laws: Understanding Transformer Performance with Associative MemoryCoin3D: Controllable and Interactive 3D Assets Generation with Proxy-Guided ConditioningHunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Unde…
…
continue reading
1
Vision-Language Model Design, Online RLHF Workflow, Multilingual AI, AI Memory Solution
9:41
9:41
Play later
Play later
Lists
Like
Liked
9:41
What matters when building vision-language models?RLHF Workflow: From Reward Modeling to Online RLHFSUTRA: Scalable Multilingual Language Model ArchitectureSambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of ExpertsPlot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from …
…
continue reading
1
BlenderAlchemy Revolution, Stylus Adapter Magic, DressCode Digital Fashion
10:04
10:04
Play later
Play later
Lists
Like
Liked
10:04
BlenderAlchemy: Editing 3D Graphics with Vision-Language ModelsStylus: Automatic Adapter Selection for Diffusion ModelsAg2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action RepresentationsDressCode: Autoregressively Sewing and Generating Garments from Text GuidancePLLaVA : Parameter-free LLaVA Extension from Images to V…
…
continue reading
1
Real-Time Motion Control, Next-Gen Visual Captions, 3D Scene Reconstruction Innovations
11:30
11:30
Play later
Play later
Lists
Like
Liked
11:30
MotionLCM: Real-time Controllable Motion Generation via Latent Consistency ModelVisual Fact Checker: Enabling High-Fidelity Detailed Caption GenerationGS-LRM: Large Reconstruction Model for 3D Gaussian SplattingSAGS: Structure-Aware 3D Gaussian SplattingInvisible Stitch: Generating Smooth 3D Scenes with Depth Inpainting…
…
continue reading
1
Kolmogorov-Arnold Networks, Iterative Reasoning Optimization, Extending Llama-3 Context Length
11:24
11:24
Play later
Play later
Lists
Like
Liked
11:24
KAN: Kolmogorov-Arnold NetworksInstantFamily: Masked Attention for Zero-shot Multi-ID Image GenerationBetter & Faster Large Language Models via Multi-token PredictionIterative Reasoning Preference OptimizationExtending Llama-3's Context Ten-Fold Overnight
…
continue reading
1
Innovative Image Editing, Advanced Autonomous Tracking, and the Evolution of Open-Source AI
12:10
12:10
Play later
Play later
Lists
Like
Liked
12:10
Paint by Inpaint: Learning to Add Image Objects by Removing Them FirstSelf-Play Preference Optimization for Language Model AlignmentAutomatic Creative Selection with Cross-Modal MatchingSTT: Stateful Tracking with Transformers for Autonomous DrivingOctopus v4: Graph of language models
…
continue reading
Adapting to poor air quality Science Sessions are brief conversations with cutting-edge researchers, National Academy members, and policymakers as they discuss topics relevant to today's scientific community. Learn the behind-the-scenes story of work published in the Proceedings of the National Academy of Sciences (PNAS), plus a broad range of scie…
…
continue reading