Large Language Models

GPT, Claude, Gemini, and other LLMs

This Week's Best | Monthly Best | Guide | Trending

Top This Week

Llms

Associative memory system for LLMs that learns during inference [P]

I've been working on MDA (Modular Dynamic Architecture), an online associative memory system for LLMs. Here's what I learned building it....

Reddit - Machine Learning · 1 min · about 3 hours ago

Llms

Things I got wrong building a confidence evaluator for local LLMs [D]

I've been building **Autodidact**, a local-first AI agent framework. The central piece is a **confidence evaluator** - something that dec...

Reddit - Machine Learning · 1 min · about 4 hours ago

Llms

I’m convinced 90% of you building "AI Agents" are just burning money on proxy providers. [D]

Seriously, I just audited my stack and realized I’m spending more on rotating residential proxies than I am on the actual Claude and Open...

Reddit - Machine Learning · 1 min · about 4 hours ago

All Content

Llms

[2512.10534] Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning

Abstract page for arXiv paper 2512.10534: Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforceme...

arXiv - AI · 4 min · about 2 months ago

Llms

[2601.22571] PerfGuard: A Performance-Aware Agent for Visual Content Generation

Abstract page for arXiv paper 2601.22571: PerfGuard: A Performance-Aware Agent for Visual Content Generation

arXiv - AI · 4 min · about 2 months ago

Llms

[2512.14106] HydroGEM: A Self Supervised Zero Shot Hybrid TCN Transformer Foundation Model for Continental Scale Streamflow Quality Control

Abstract page for arXiv paper 2512.14106: HydroGEM: A Self Supervised Zero Shot Hybrid TCN Transformer Foundation Model for Continental S...

arXiv - AI · 4 min · about 2 months ago

Llms

[2512.07081] ClinNoteAgents: An LLM Multi-Agent System for Predicting and Interpreting Heart Failure 30-Day Readmission from Clinical Notes

Abstract page for arXiv paper 2512.07081: ClinNoteAgents: An LLM Multi-Agent System for Predicting and Interpreting Heart Failure 30-Day ...

arXiv - AI · 4 min · about 2 months ago

Llms

[2505.13770] Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference

Abstract page for arXiv paper 2505.13770: Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Infe...

arXiv - Machine Learning · 4 min · about 2 months ago

Llms

[2511.21033] Towards Trustworthy Legal AI through LLM Agents and Formal Reasoning

Abstract page for arXiv paper 2511.21033: Towards Trustworthy Legal AI through LLM Agents and Formal Reasoning

arXiv - AI · 4 min · about 2 months ago

Llms

[2511.04439] CoRPO: Adding a Correctness Bias to GRPO Improves Generalization

Abstract page for arXiv paper 2511.04439: CoRPO: Adding a Correctness Bias to GRPO Improves Generalization

arXiv - Machine Learning · 4 min · about 2 months ago

Llms

[2510.08966] Beyond Prefixes: Graph-as-Memory Cross-Attention for Knowledge Graph Completion with Large Language Models

Abstract page for arXiv paper 2510.08966: Beyond Prefixes: Graph-as-Memory Cross-Attention for Knowledge Graph Completion with Large Lang...

arXiv - AI · 4 min · about 2 months ago

Llms

[2505.04997] Foam-Agent: Towards Automated Intelligent CFD Workflows

Abstract page for arXiv paper 2505.04997: Foam-Agent: Towards Automated Intelligent CFD Workflows

arXiv - AI · 3 min · about 2 months ago

Llms

[2503.07928] The StudyChat Dataset: Analyzing Student Dialogues With ChatGPT in an Artificial Intelligence Course

Abstract page for arXiv paper 2503.07928: The StudyChat Dataset: Analyzing Student Dialogues With ChatGPT in an Artificial Intelligence C...

arXiv - AI · 4 min · about 2 months ago

Llms

[2603.05500] POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation

Abstract page for arXiv paper 2603.05500: POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation

arXiv - Machine Learning · 3 min · about 2 months ago

Llms

[2603.05494] Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation

Abstract page for arXiv paper 2603.05494: Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation

arXiv - Machine Learning · 4 min · about 2 months ago

Llms

[2603.05488] Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought

Abstract page for arXiv paper 2603.05488: Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought

arXiv - Machine Learning · 3 min · about 2 months ago

Llms

[2603.05471] Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval

Abstract page for arXiv paper 2603.05471: Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval

arXiv - AI · 4 min · about 2 months ago

Llms

[2603.05432] Ensembling Language Models with Sequential Monte Carlo

Abstract page for arXiv paper 2603.05432: Ensembling Language Models with Sequential Monte Carlo

arXiv - Machine Learning · 4 min · about 2 months ago

Llms

[2603.05421] MobileFetalCLIP: Selective Repulsive Knowledge Distillation for Mobile Fetal Ultrasound Analysis

Abstract page for arXiv paper 2603.05421: MobileFetalCLIP: Selective Repulsive Knowledge Distillation for Mobile Fetal Ultrasound Analysis

arXiv - Machine Learning · 3 min · about 2 months ago

Llms

[2603.05308] Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution

Abstract page for arXiv paper 2603.05308: Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution

arXiv - AI · 4 min · about 2 months ago

Llms

[2603.05210] Balancing Coverage and Draft Latency in Vocabulary Trimming for Faster Speculative Decoding

Abstract page for arXiv paper 2603.05210: Balancing Coverage and Draft Latency in Vocabulary Trimming for Faster Speculative Decoding

arXiv - Machine Learning · 4 min · about 2 months ago

Llms

[2603.05299] WavSLM: Single-Stream Speech Language Modeling via WavLM Distillation

Abstract page for arXiv paper 2603.05299: WavSLM: Single-Stream Speech Language Modeling via WavLM Distillation

arXiv - Machine Learning · 3 min · about 2 months ago

Llms

[2603.05167] C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning

Abstract page for arXiv paper 2603.05167: C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reas...

arXiv - AI · 3 min · about 2 months ago

Previous Page 239 Next

Stay updated with AI News

Get the latest news, tools, and insights delivered to your inbox.

Subscribe to Newsletter

Daily or weekly digest • Unsubscribe anytime

Large Language Models

Top This Week

Associative memory system for LLMs that learns during inference [P]

Things I got wrong building a confidence evaluator for local LLMs [D]

I’m convinced 90% of you building "AI Agents" are just burning money on proxy providers. [D]

All Content

[2512.10534] Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning

[2601.22571] PerfGuard: A Performance-Aware Agent for Visual Content Generation

[2512.14106] HydroGEM: A Self Supervised Zero Shot Hybrid TCN Transformer Foundation Model for Continental Scale Streamflow Quality Control

[2512.07081] ClinNoteAgents: An LLM Multi-Agent System for Predicting and Interpreting Heart Failure 30-Day Readmission from Clinical Notes

[2505.13770] Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference

[2511.21033] Towards Trustworthy Legal AI through LLM Agents and Formal Reasoning

[2511.04439] CoRPO: Adding a Correctness Bias to GRPO Improves Generalization

[2510.08966] Beyond Prefixes: Graph-as-Memory Cross-Attention for Knowledge Graph Completion with Large Language Models

[2505.04997] Foam-Agent: Towards Automated Intelligent CFD Workflows

[2503.07928] The StudyChat Dataset: Analyzing Student Dialogues With ChatGPT in an Artificial Intelligence Course

[2603.05500] POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation

[2603.05494] Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation

[2603.05488] Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought

[2603.05471] Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval

[2603.05432] Ensembling Language Models with Sequential Monte Carlo

[2603.05421] MobileFetalCLIP: Selective Repulsive Knowledge Distillation for Mobile Fetal Ultrasound Analysis

[2603.05308] Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution

[2603.05210] Balancing Coverage and Draft Latency in Vocabulary Trimming for Faster Speculative Decoding

[2603.05299] WavSLM: Single-Stream Speech Language Modeling via WavLM Distillation

[2603.05167] C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning

Related Topics

Stay updated with AI News