mirror of
https://github.com/Shopify/liquid.git
synced 2026-09-15 08:50:45 -07:00
To obtain String Literal, the regular expression might be executed at two times.
It would be slightly faster to run them all at once using `Regexp.union`.
− | before | after | result
-- | -- | -- | --
parse | 63.418 | 65.183 | 1.028x
render | 195.389 | 195.648 | -
parse & render | 46.091 | 46.917 | 1.018x
### Environment
- MacBook Pro (14 inch, 2021)
- macOS 13.0 Beta
- Apple M1 Max
- Ruby 3.1.2
### Before
```
Running benchmark for 10 seconds (with 5 seconds warmup).
Warming up --------------------------------------
parse: 6.000 i/100ms
render: 19.000 i/100ms
parse & render: 4.000 i/100ms
Calculating -------------------------------------
parse: 63.418 (± 0.0%) i/s - 636.000 in 10.028939s
render: 195.389 (± 0.5%) i/s - 1.957k in 10.016466s
parse & render: 46.091 (± 0.0%) i/s - 464.000 in 10.067445s
```
### After
```
Running benchmark for 10 seconds (with 5 seconds warmup).
Warming up --------------------------------------
parse: 6.000 i/100ms
render: 19.000 i/100ms
parse & render: 4.000 i/100ms
Calculating -------------------------------------
parse: 65.183 (± 0.0%) i/s - 654.000 in 10.033549s
render: 195.648 (± 1.0%) i/s - 1.957k in 10.003511s
parse & render: 46.917 (± 0.0%) i/s - 472.000 in 10.060782s
```
62 lines
1.5 KiB
Ruby
62 lines
1.5 KiB
Ruby
# frozen_string_literal: true
|
|
|
|
require "strscan"
|
|
module Liquid
|
|
class Lexer
|
|
SPECIALS = {
|
|
'|' => :pipe,
|
|
'.' => :dot,
|
|
':' => :colon,
|
|
',' => :comma,
|
|
'[' => :open_square,
|
|
']' => :close_square,
|
|
'(' => :open_round,
|
|
')' => :close_round,
|
|
'?' => :question,
|
|
'-' => :dash,
|
|
}.freeze
|
|
IDENTIFIER = /[a-zA-Z_][\w-]*\??/
|
|
SINGLE_STRING_LITERAL = /'[^\']*'/
|
|
DOUBLE_STRING_LITERAL = /"[^\"]*"/
|
|
STRING_LITERAL = Regexp.union(SINGLE_STRING_LITERAL, DOUBLE_STRING_LITERAL)
|
|
NUMBER_LITERAL = /-?\d+(\.\d+)?/
|
|
DOTDOT = /\.\./
|
|
COMPARISON_OPERATOR = /==|!=|<>|<=?|>=?|contains(?=\s)/
|
|
WHITESPACE_OR_NOTHING = /\s*/
|
|
|
|
def initialize(input)
|
|
@ss = StringScanner.new(input)
|
|
end
|
|
|
|
def tokenize
|
|
@output = []
|
|
|
|
until @ss.eos?
|
|
@ss.skip(WHITESPACE_OR_NOTHING)
|
|
break if @ss.eos?
|
|
tok = if (t = @ss.scan(COMPARISON_OPERATOR))
|
|
[:comparison, t]
|
|
elsif (t = @ss.scan(STRING_LITERAL))
|
|
[:string, t]
|
|
elsif (t = @ss.scan(NUMBER_LITERAL))
|
|
[:number, t]
|
|
elsif (t = @ss.scan(IDENTIFIER))
|
|
[:id, t]
|
|
elsif (t = @ss.scan(DOTDOT))
|
|
[:dotdot, t]
|
|
else
|
|
c = @ss.getch
|
|
if (s = SPECIALS[c])
|
|
[s, c]
|
|
else
|
|
raise SyntaxError, "Unexpected character #{c}"
|
|
end
|
|
end
|
|
@output << tok
|
|
end
|
|
|
|
@output << [:end_of_string]
|
|
end
|
|
end
|
|
end
|