Search

Tag: #apple-ml13 results

Apple's Async Verified Semantic Caching for LLMs

Apple's Async Verified Semantic Caching for LLMs

Apple introduces asynchronous verified semantic caching to optimize tiered LLM architectures. It addresses tradeoffs in static and dynamic caches using embedding similarity thresholds. This reduces inference cost and latency in production workflows like search and agents.

Apple Machine LearningOfficialFeb 16#research#apple-ml#llm
Hyperparameter Transfer Across All Scaling Axes

Hyperparameter Transfer Across All Scaling Axes

Apple ML extends μP for hyperparameter transfer across model sizes, modules, width, depth, batch, and duration. Introduces Complete(d) Parameterisation unifying width-depth scaling. Enables optimal base hyperparameters search at small scales for large model transfer.

Apple Machine LearningOfficialFeb 13#research#apple-ml#mu-p
Hyperparam Transfer Across All Scales

Hyperparam Transfer Across All Scales

Apple extends μP for hyperparameter transfer across modules, width, depth, batch, and duration. Introduces Complete(d) Parameterisation unifying width-depth scaling. Enables optimal hypers from small to large models.

Apple Machine LearningOfficialFeb 13#research#apple-ml#mu-p
Faster Rates for Federated VIs

Faster Rates for Federated VIs

Apple advances federated optimization for stochastic variational inequalities. Establishes improved convergence rates closing gap with convex optimization. Refined analysis boosts Local Extra SGD for smooth monotone VIs.

Apple Machine LearningOfficialFeb 13#research#apple-ml#local-extra-sgd
Complete Hyperparameter Transfer for Scaling

Complete Hyperparameter Transfer for Scaling

Apple ML extends μP parameterisations with Complete(d) Parameterisation for hyperparameter transfer. Covers scaling across modules, width, depth, batch size, and duration. Enables optimal hyperparameter search on small models for transfer to large-scale ones.

Apple Machine LearningOfficialFeb 13#research#apple-ml#mu-p
Cadmus: Low-Cost Program Synthesis System

Cadmus: Low-Cost Program Synthesis System

Apple ML introduces Cadmus, a small-scale system for autoregressive program synthesis. It features an integer virtual machine, a dataset of diverse true programs, and a transformer model trained for under $200 compute. This setup enables controlled experiments bypassing issues with large LLMs like OOD challenges and high resource demands.

Apple Machine LearningOfficialFeb 13#research#apple-ml#cadmus
Cadmus: Cheap Program Synthesis System

Cadmus: Cheap Program Synthesis System

Apple unveils Cadmus, a small-scale system for autoregressive program synthesis. It features an integer VM, diverse program dataset, and transformer model trained under $200 compute. Enables controlled experiments bypassing LLM challenges like OOD and tokenization.

Apple Machine LearningOfficialFeb 13#research#apple-ml#cadmus
Cadmus: Affordable Autoregressive Program Synthesis

Cadmus: Affordable Autoregressive Program Synthesis

Apple ML introduces Cadmus, a small-scale system for autoregressive program synthesis. It features an integer virtual machine, a dataset of diverse true programs, and a transformer model trained for under $200 compute. This setup enables controlled experiments avoiding LLM pitfalls like OOD issues and high compute demands.

Apple Machine LearningOfficialFeb 13#research#apple-ml#cadmus
Page 1 of 2