Files
liquid/OPTIMIZATION.md
Tobi Lutke c5a44be104 perf: optimize compiled template allocations and performance
Major optimizations to reduce allocations and improve execution speed:

1. For loops: Replace catch/throw with while + break flag
   - Uses while loop with index instead of .each with catch/throw
   - Break implemented with flag variable, continue with next
   - Result: 18% fewer allocations, 85% faster for simple loops

2. Forloop property inlining
   - Inline forloop.index as (__idx__ + 1), forloop.first as (__idx__ == 0), etc.
   - Completely eliminates forloop hash allocation when all properties inlinable
   - Result: Loop with forloop went from +46% MORE to -16% FEWER allocations

3. LR.to_array helper with EMPTY_ARRAY constant
   - Centralized array conversion with frozen empty array for nil
   - Avoids allocations for empty collections

4. Inline LR.truthy? calls
   - Replace LR.truthy?(x) with (x != nil && x != false)
   - Eliminates method call overhead in conditions

5. Keep Time methods available in sandbox for date filter

Overall results:
- Allocations: 3.5% MORE -> 24% FEWER (27% improvement)
- Time: 64% faster -> 89% faster (25% improvement)

Also adds:
- compile_profiler.rb for measuring allocations/performance
- compile_acceptance_test.rb for output equivalence testing
- OPTIMIZATION.md documenting optimization status
2025-12-31 13:24:33 -04:00

2.1 KiB

Liquid Compiled Template Optimization Log

This document tracks optimizations made to the compiled Liquid template engine. Each entry shows before/after code and measured impact.


Baseline Measurement

Date: 2024-12-31 Commit: (pending profiler implementation)

Current State

The compiled template engine generates Ruby code from Liquid templates. Before optimizations, here's a sample of generated code for a simple loop:

# Template: {% for product in products %}{{ forloop.index }}: {{ product.name }}{% endfor %}

->(assigns, __context__, __external__) do
  __output__ = +""
  
  __coll1__ = assigns["products"]
  __coll1__ = __coll1__.to_a if __coll1__.is_a?(Range)
  __len3__ = __coll1__.respond_to?(:length) ? __coll1__.length : 0
  __idx2__ = 0
  catch(:__loop__break__) do
    (__coll1__.respond_to?(:each) ? __coll1__ : []).each do |__item__|
      catch(:__loop__continue__) do
        assigns["product"] = __item__
        assigns['forloop'] = {
          'name' => "product-products",
          'length' => __len3__,
          'index' => __idx2__ + 1,
          'index0' => __idx2__,
          'rindex' => __len3__ - __idx2__,
          'rindex0' => __len3__ - __idx2__ - 1,
          'first' => __idx2__ == 0,
          'last' => __idx2__ == __len3__ - 1,
        }
        __output__ << LR.output(LR.lookup(assigns["forloop"], "index", __context__))
        __output__ << ": "
        __output__ << LR.output(LR.lookup(assigns["product"], "name", __context__))
      end
      __idx2__ += 1
    end
  end
  assigns.delete("product")
  assigns.delete('forloop')
  
  __output__
end

Issues Identified

  1. catch/throw overhead - Used even when no break/continue in loop
  2. Hash allocation per iteration - 8 key/value pairs computed every time
  3. respond_to? checks - Redundant after type is known
  4. LR.lookup for forloop - Unnecessary indirection for known hash
  5. String literals not frozen - Allocates on each render
  6. Output buffer grows dynamically - No pre-allocation

Optimization Log