Repository navigation
Optimisations - #16
Closed
ankushv-003 wants to merge 1 commit into
Closed
Optimisations #16ankushv-003 wants to merge 1 commit into
ankushv-003 wants to merge 1 commit into
Conversation
Owner
|
Closing AI slop PR. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
go-polyline Optimization Report
Executive Summary
This document details all optimizations implemented and potential future optimizations for the
go-polylinepackage.Implemented Optimizations
1. Unrolled DecodeUint Fast Paths
File:
polyline.go:46-95Problem: The original loop-based decoder had branch overhead for every byte, even though ~95% of coordinate deltas encode to 1-4 bytes.
Solution: Unrolled decoding for 1-4 byte values without loop overhead.
Results:
2. Unrolled EncodeUint Fast Paths
File:
polyline.go:148-168Problem: Small deltas required loop iterations even for simple cases.
Solution: Direct encoding for values < 1,048,576 (covers 99%+ of cases).
Results:
3. Specialized Dim=2 Decoders
File:
polyline_fast.go:6-46Problem: Generic N-dimensional decoder had:
make([]int, c.Dim)allocation per callSolution: Dedicated decode path for 2D coordinates with inlined lat/lon handling.
Results (DecodeCoords):
4. Specialized Dim=2 Encoders with Pre-allocation
File:
polyline_fast.go:75-137Problem:
append(), causing multiple reallocationsSolution:
Results (EncodeCoords):
5. Streaming APIs (Zero-Allocation)
File:
stream.goProblem: Batch APIs allocate a new
[]float64slice for every coordinate, even when the caller processes them sequentially.Solution: New streaming
Decoder2DandEncoder2Dtypes that reuse internal buffers.Results:
6. Bounds Check Elimination (BCE) Hints
File:
polyline_fast.go:92, 124Problem: Go compiler cannot always prove array bounds are safe, generating redundant checks.
Solution: Explicit bounds access before loops helps BCE optimization pass.
7. Simplified round() Function
File:
polyline.go:32-34Problem: Manual rounding implementation had branches.
Solution: Use
math.Round()which is compiler-optimized.Potential Future Optimizations
1. Fixed-Size Coordinate Type (
[2]float64)Status: NOT IMPLEMENTED (API breaking change)
Problem: Current
DecodeCoordsreturns[][]float64, allocating a new slice for each coordinate. This is the #1 allocation source.Solution: Return
[][2]float64or[]Coord2Dwheretype Coord2D [2]float64.Benchmark Results:
Memory Usage:
Why Not Implemented: This would be a breaking API change. Users would need to migrate from
[][]float64to[]Coord2D. Consider adding as a new API alongside existing one.2. Fully Inlined Integer Decoding
Status: NOT IMPLEMENTED (code complexity vs. benefit tradeoff)
Problem: Function call overhead for
DecodeIntin tight loops.Solution: Inline the entire decode logic including zigzag decoding.
Benchmark Results:
Combined with Array Type:
Why Not Implemented:
3. Contiguous Backing Array
Status: NOT IMPLEMENTED (marginal benefit)
Problem: Even with
[][]float64, each inner slice has overhead.Solution: Single
[]float64backing array with slice headers pointing into it.Benchmark Results:
Why Not Implemented:
4. SIMD/AVX Vectorization
Status: NOT FEASIBLE
Problem: Could theoretically decode multiple bytes in parallel.
Why Not Feasible: Polyline encoding uses variable-length integers. You cannot know where value N+1 starts until you fully decode value N. This sequential dependency fundamentally prevents SIMD parallelization.
5. sync.Pool for High-Concurrency Scenarios
Status: NOT IMPLEMENTED (streaming APIs solve this better)
Problem: In high-concurrency scenarios, frequent allocations cause GC pressure.
Solution: Pool coordinate slices for reuse.
Why Not Implemented:
Decoder2D,Encoder2D) already achieve zero allocationsRecommended API Additions
To unlock the remaining performance gains without breaking existing APIs, consider adding:
This provides:
Summary of All Optimizations
Implemented (Current State)
Potential (Not Implemented)
[2]float64typeThroughput Summary
Coord2D(potential)