Per-program id timings #17554

sakridge · 2021-05-27T19:15:07Z

Problem

Hard to tell which program is taking the longest in the network.

Summary of Changes

Add per-program timings configurable with a validator flag.

Fixes #

codecov · 2021-06-01T15:11:58Z

Codecov Report

Merging #17554 (bbd1da5) into master (bbcdf07) will decrease coverage by 0.0%.
The diff coverage is 90.3%.

@@            Coverage Diff            @@
##           master   #17554     +/-   ##
=========================================
- Coverage    82.7%    82.7%   -0.1%     
=========================================
  Files         430      430             
  Lines      120566   120594     +28     
=========================================
- Hits        99820    99818      -2     
- Misses      20746    20776     +30

core/src/progress_map.rs

core/src/validator.rs

tao-stones · 2021-06-03T16:31:19Z

runtime/src/message_processor.rs

+            time.stop();
+            if timings.collect_per_program_timings {
+                let program_id = instruction.program_id(&message.account_keys);
+                *timings.per_program_timings.entry(*program_id).or_insert(0) += time.as_us();


In the case of same program being executed multiple times in a transaction, do we want to see the accumulated execution time in stats report, or perhaps averaged time?
Cost model interested in later, if reporting accumulated time is desired, it is always possible to take average at the time of updating cost model.

yes. I think we'll have to have a count stat here as well. I'll add that.

tao-stones

looks good

* Cost Model to limit transactions which are not parallelizeable (#16694) * * Add following to banking_stage: 1. CostModel as immutable ref shared between threads, to provide estimated cost for transactions. 2. CostTracker which is shared between threads, tracks transaction costs for each block. * replace hard coded program ID with id() calls * Add Account Access Cost as part of TransactionCost. Account Access cost are weighted differently between read and write, signed and non-signed. * Establish instruction_execution_cost_table, add function to update or insert instruction cost, unit tested. It is read-only for now; it allows Replay to insert realtime instruction execution costs to the table. * add test for cost_tracker atomically try_add operation, serves as safety guard for future changes * check cost against local copy of cost_tracker, return transactions that would exceed limit as unprocessed transaction to be buffered; only apply bank processed transactions cost to tracker; * bencher to new banking_stage with max cost limit to allow cost model being hit consistently during bench iterations * replay stage feed back program cost (#17731) * replay stage feeds back realtime per-program execution cost to cost model; * program cost execution table is initialized into empty table, no longer populated with hardcoded numbers; * changed cost unit to microsecond, using value collected from mainnet; * add ExecuteCostTable with fixed capacity for security concern, when its limit is reached, programs with old age AND less occurrence will be pushed out to make room for new programs. * investigate system performance test degradation (#17919) * Add stats and counter around cost model ops, mainly: - calculate transaction cost - check transaction can fit in a block - update block cost tracker after transactions are added to block - replay_stage to update/insert execution cost to table * Change mutex on cost_tracker to RwLock * removed cloning cost_tracker for local use, as the metrics show clone is very expensive. * acquire and hold locks for block of TXs, instead of acquire and release per transaction; * remove redundant would_fit check from cost_tracker update execution path * refactor cost checking with less frequent lock acquiring * avoid many Transaction_cost heap allocation when calculate cost, which is in the hot path - executed per transaction. * create hashmap with new_capacity to reduce runtime heap realloc. * code review changes: categorize stats, replace explicit drop calls, concisely initiate to default * address potential deadlock by acquiring locks one at time * Persist cost table to blockstore (#18123) * Add `ProgramCosts` Column Family to blockstore, implement LedgerColumn; add `delete_cf` to Rocks * Add ProgramCosts to compaction excluding list alone side with TransactionStatusIndex in one place: `excludes_from_compaction()` * Write cost table to blockstore after `replay_stage` replayed active banks; add stats to measure persist time * Deletes program from `ProgramCosts` in blockstore when they are removed from cost_table in memory * Only try to persist to blockstore when cost_table is changed. * Restore cost table during validator startup * Offload `cost_model` related operations from replay main thread to dedicated service thread, add channel to send execute_timings between these threads; * Move `cost_update_service` to its own module; replay_stage is now decoupled from cost_model. * log warning when channel send fails (#18391) * Aggregate cost_model into cost_tracker (#18374) * * aggregate cost_model into cost_tracker, decouple it from banking_stage to prevent accidental deadlock. * Simplified code, removed unused functions * review fixes * update ledger tool to restore cost table from blockstore (#18489) * update ledger tool to restore cost model from blockstore when compute-slot-cost * Move initialize_cost_table into cost_model, so the function can be tested and shared between validator and ledger-tool * refactor and simplify a test * manually fix merge conflicts * Per-program id timings (#17554) * more manual fixing * solve a merge conflict * featurize cost model * more merge fix * cost model uses compute_unit to replace microsecond as cost unit (#18934) * Reject blocks for costs above the max block cost (#18994) * Update block max cost limit to fix performance regession (#19276) * replace function with const var for better readability (#19285) * Add few more metrics data points (#19624) * periodically report sigverify_stage stats (#19674) * manual merge * cost model nits (#18528) * Accumulate consumed units (#18714) * tx wide compute budget (#18631) * more manual merge * ignore zerorize drop security * - update const cost values with data collected by #19627 - update cost calculation to closely proposed fee schedule #16984 * add transaction cost histogram metrics (#20350) * rebase to 1.7.15 * add tx count and thread id to stats (#20451) each stat reports and resets when slot changes * remove cost_model feature_set * ignore vote transactions from cost model Co-authored-by: sakridge <[email protected]> Co-authored-by: Jeff Biseda <[email protected]> Co-authored-by: Jack May <[email protected]>

sakridge force-pushed the per-program-id-timings branch 8 times, most recently from 5c2a9f2 to 020ab4e Compare June 1, 2021 13:41

sakridge mentioned this pull request Jun 3, 2021

Replay Stage feeds realtime instruction execution cost to Cost Model #17696

Closed

tao-stones reviewed Jun 3, 2021

View reviewed changes

core/src/progress_map.rs Show resolved Hide resolved

core/src/validator.rs Outdated Show resolved Hide resolved

tao-stones reviewed Jun 3, 2021

View reviewed changes

sakridge force-pushed the per-program-id-timings branch 2 times, most recently from 33911c4 to 3a195da Compare June 3, 2021 19:01

Per-program id timings

bbd1da5

sakridge force-pushed the per-program-id-timings branch from 3a195da to bbd1da5 Compare June 3, 2021 19:05

tao-stones approved these changes Jun 3, 2021

View reviewed changes

sakridge merged commit f97ce2c into solana-labs:master Jun 4, 2021

sakridge deleted the per-program-id-timings branch June 4, 2021 14:04

tao-stones pushed a commit to tao-stones/solana that referenced this pull request Jul 16, 2021

Per-program id timings (solana-labs#17554)

ca6a7b2

tao-stones pushed a commit to tao-stones/solana that referenced this pull request Jul 16, 2021

Per-program id timings (solana-labs#17554)

6e4fe4a

tao-stones pushed a commit to tao-stones/solana that referenced this pull request Jul 17, 2021

Per-program id timings (solana-labs#17554)

0b616d2

tao-stones pushed a commit to tao-stones/solana that referenced this pull request Sep 24, 2021

Per-program id timings (solana-labs#17554)

a715df7

tao-stones pushed a commit to tao-stones/solana that referenced this pull request Sep 29, 2021

Per-program id timings (solana-labs#17554)

e11d1a0

tao-stones pushed a commit to tao-stones/solana that referenced this pull request Sep 30, 2021

Per-program id timings (solana-labs#17554)

c5e5b5a

tao-stones pushed a commit to tao-stones/solana that referenced this pull request Oct 6, 2021

Per-program id timings (solana-labs#17554)

731d1bc

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Per-program id timings #17554

Per-program id timings #17554

sakridge commented May 27, 2021

codecov bot commented Jun 1, 2021 •

edited

Loading

tao-stones Jun 3, 2021

sakridge Jun 3, 2021

tao-stones left a comment

Per-program id timings #17554

Per-program id timings #17554

Conversation

sakridge commented May 27, 2021

Problem

Summary of Changes

codecov bot commented Jun 1, 2021 • edited Loading

Codecov Report

tao-stones Jun 3, 2021

Choose a reason for hiding this comment

sakridge Jun 3, 2021

Choose a reason for hiding this comment

tao-stones left a comment

Choose a reason for hiding this comment

codecov bot commented Jun 1, 2021 •

edited

Loading