Posts

AutoRecon: Automated 3D Object Discovery and Reconstruction #1107

Image
AutoRecon: Automated 3D Object Discovery and Reconstruction #1107 Source  

Make-An-Animation: Large-Scale Text-conditional 3D Human Motion Generation #1106

Image
Make-An-Animation: Large-Scale Text-conditional 3D Human Motion Generation #1106 Source  

Understanding 3D Object Interaction from a Single Image #1105

Image
Understanding 3D Object Interaction from a Single Image #1105 Source  

FitMe: Deep Photorealistic 3D Morphable Model Avatars #1104

Image
FitMe: Deep Photorealistic 3D Morphable Model Avatars #1104 Source

Improving Language Model Negotiation with Self-Play and In-Context Learning from AI Feedback #1103

Image
Improving Language Model Negotiation with Self-Play and In-Context Learning from AI Feedback #1103 Source  

Google's SoundStorm: Efficient Parallel Audio Generation - Produces audio of comparable quality to AR models while being two orders of magnitude faster #1102

Image
Google's SoundStorm: Efficient Parallel Audio Generation - Produces audio of comparable quality to AR models while being two orders of magnitude faster #1102 Source  

Introducing Phoenix: a revolutionary humanoid general-purpose robot designed for work #1101

Image
Introducing Phoenix: a revolutionary humanoid general-purpose robot designed for work #1101 Source  

Multiple fully Tesla-made Bots now walking around & learning about the real world #1100

Image
Multiple fully Tesla-made Bots now walking around & learning about the real world #1100 Source  

Google: Using reinforcement learning for dynamic planning in open-ended conversations #1099

Image
Google: Using reinforcement learning for dynamic planning in open-ended conversations #1099 Source  

Microsoft's TinyStories: How Small Can Language Models Be and Still Speak Coherent English #1098

Image
Microsoft's TinyStories: How Small Can Language Models Be and Still Speak Coherent English #1098 Source  

Consensus and subjectivity of skin tone annotation for ML fairness #1097

Image
Consensus and subjectivity of skin tone annotation for ML fairness #1097 Source  

HACK: Learning a Parametric Head and Neck Model for High-fidelity Animation #1096

Image
HACK: Learning a Parametric Head and Neck Model for High-fidelity Animation #1096 Source  

MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers #1095

Image
MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers #1095 Source  

100k context windows now available on Poe: we are excited to start a beta test for Claude-instant-100k #1094

Image
100k context windows now available on Poe: we are excited to start a beta test for Claude-instant-100k #1094 Source  

OpenAI: We’re rolling out web browsing and Plugins to all ChatGPT Plus users over the next week! Moving from alpha to beta, they allow ChatGPT to access the internet and to use 70+ third-party Plugins #1093

Image
OpenAI: We’re rolling out web browsing and Plugins to all ChatGPT Plus users over the next week! Moving from alpha to beta, they allow ChatGPT to access the internet and to use 70+ third-party Plugins #1093 Source  

Artificial intelligence identifies anti-aging drug candidates targeting 'zombie' cells #1092

Image
Artificial intelligence identifies anti-aging drug candidates targeting 'zombie' cells #1092 Source  

Assisted Generation: a new direction toward low-latency text generation (3x faster) #1091

Image
Assisted Generation: a new direction toward low-latency text generation (3x faster) Source  

EfficientViT: Memory Efficient Vision Transformer with Cascaded Group Attention #1090

Image
EfficientViT: Memory Efficient Vision Transformer with Cascaded Group Attention #1090 Source  

AnthropicAI expanded Claude’s context window to 100,000 tokens of text, corresponding to around 75K words #1089

Image
AnthropicAI expanded Claude’s context window to 100,000 tokens of text, corresponding to around 75K words #1089 Source  

Synthesia research releases HumanRF: High-Fidelity Neural Radiance Fields for Humans in Motion #1088

Image
Synthesia research releases HumanRF: High-Fidelity Neural Radiance Fields for Humans in Motion #1088 Source  

Google IO 2023 wrapped up #1087

Image
Google IO 2023 wrapped up #1087 Source  

Google's Generative AI is coming to search #1086

Image
Google's Generative AI is coming to search #1086 Source  

Google introduces PaLM 2 and it is coming to more than 25 products of Google #1085

Image
Google introduces PaLM 2 and it is coming to more than 25 products of Google #1085 Source  

Bard available in over 180 countries and territories including India and upgraded to PaLM 2 #1084

Image
Bard available in over 180 countries and territories including India and upgraded to PaLM 2 #1084 Source  

Hugging Face's Transformer Agents - Control 100,000+ HF models by talking to Transformers and Diffusers #1083

Image
Hugging Face's Transformer Agents - Control 100,000+ HF models by talking to Transformers and Diffusers #1083 Source  

Announcing the Haven-1 and Vast-1 missions to low-Earth orbit. Launched by SpaceX, Haven-1 is scheduled to be the world’s first commercial space station #1082

Image
Announcing the Haven-1 and Vast-1 missions to low-Earth orbit. Launched by SpaceX, Haven-1 is scheduled to be the world’s first commercial space station #1082 Source  

MultiModal-GPT: A Vision and Language Model for Dialogue with Humans #1081

Image
MultiModal-GPT: A Vision and Language Model for Dialogue with Humans #1081 Source  

FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance #1080

Image
FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance #1080 Source  

TidyBot: Personalized Robot Assistance with Large Language Models #1079

Image
TidyBot: Personalized Robot Assistance with Large Language Models #1079 Source  

OpenAI GPT-4 to interpretability — automatically proposing explanations for GPT-2's 300k Neurons #1078

Image
OpenAI GPT-4 to interpretability — automatically proposing explanations for GPT-2's 300k Neurons #1078 Source  

Introducing LeMUR, short for Leveraging Large Language Models to Understand Recognized Speech #1077

Image
Introducing LeMUR, short for Leveraging Large Language Models to Understand Recognized Speech #1077 Source  

ImageBind: Meta's latest multimodal embedding, covering not only the usual suspects (text, image, audio), but also depth, thermal (infrared), and IMU signals (open source) #1076

Image
ImageBind: Meta's latest multimodal embedding, covering not only the usual suspects (text, image, audio), but also depth, thermal (infrared), and IMU signals (open source) #1076 Source ak

Announcing Nyric, an AI world-generation platform for digital communities.Build the world of your dreams in seconds #1075

Image
Announcing Nyric, an AI world-generation platform for digital communities.Build the world of your dreams in seconds #1075 Source  

Locally Attentional SDF Diffusion for Controllable 3D Shape Generation #1074

Image
Locally Attentional SDF Diffusion for Controllable 3D Shape Generation #1074 Source  

Multi-Space Neural Radiance Fields #1073

Image
Multi-Space Neural Radiance Fields #1073 Source  

Composite Motion Learning with Task Control #1072

Image
Composite Motion Learning with Task Control #1072 Source

Dolphin is a chatbot that can interact with videos, spanning from video understanding to generation/Editing #1071

Image
Dolphin is a chatbot that can interact with videos, spanning from video understanding to generation/Editing #1071 Source  

The first RedPajama models are here! The 3B and 7B models are now available under Apache 2.0 license, including instruction-tuned and chat versions #1070

Image
The first RedPajama models are here! The 3B and 7B models are now available under Apache 2.0 license, including instruction-tuned and chat versions #1070 Source

MPT-7B: A New Standard for Open-Source, Commercially Usable LLMs (up to 65k tokens context length & trained on 1T tokens) #1069

Image
MPT-7B: A New Standard for Open-Source, Commercially Usable LLMs (up to 65k tokens context length & trained on 1T tokens) #1069 Source  

Personalize Segment Anything Model with One Shot #1068

Image
Personalize Segment Anything Model with One Shot #1068 Source  

Nvidia Real-Time Neural Appearance Models #1067

Image
Nvidia Real-Time Neural Appearance Models #1067 Source

AI Music created with Grimes voice after she open source it #1066

Image
AI Music created with Grimes voice after she open source it #1068 Kito Ether  

AutoML-GPT: Automatic Machine Learning with GPT #1065

Image
AutoML-GPT: Automatic Machine Learning with GPT #1065 Source

MaMMUT: A simple vision-encoder text-decoder architecture for multimodal tasks #1064

Image
MaMMUT: A simple vision-encoder text-decoder architecture for multimodal tasks #1064 Source

StarCoder is a 15B LLM for code with 8k context and trained only on permissive data in 80+ programming languages (open source) #1063

Image
StarCoder is a 15B LLM for code with 8k context and trained only on permissive data in 80+ programming languages (open source) #1063 Source

OpenAI's Shap-E: Generating Conditional 3D Implicit Functions #1062

Image
OpenAI's Shap-E: Generating Conditional 3D Implicit Functions #1062 Source

Walking through different worlds - In future we will have AI + AR filters that will change how our world looks & more! #1061

Image
Walking through different worlds - In future we will have AI + AR filters that will change how our world looks & more! #1061 Source  

CLIP ViT-L/14 model with 79.2% zero-shot accuracy on ImageNet #1060

Image
CLIP ViT-L/14 model with 79.2% zero-shot accuracy on ImageNet #1060 Source