Skip to content

2609 [BE] new user analytics functionality using database table - #2687

Open
andrsam wants to merge 8 commits into
masterfrom
2609
Open

andrsam wants to merge 8 commits into
masterfrom
2609

Conversation

@andrsam

@andrsam andrsam commented Mar 30, 2025 •

Copy link
Copy Markdown
Collaborator

##2609

filling analytics table using job and view analytics using data from user_analytics table
feature flags: brn.user.analytics.use.new.version - turns off using new analytics functionality
brn.user.analytics.job.enabled - enables filling database table using job

@andrsam
andrsam requested a review from ElenaSpb as a code owner March 30, 2025 07:19
@github-actions

github-actions Bot commented Mar 30, 2025 •

Copy link
Copy Markdown

Gradle Unit and Integration Test Results

555 tests  +4   554 ✔️ +4   29s ⏱️ -10s
120 suites +2       1 💤 ±0 
120 files   +2       0 ❌ ±0 

Results for commit 636d316. ± Comparison against base commit 638e332.

This pull request removes 13 and adds 17 tests. Note that renamed tests count towards both.
com.epam.brn.integration.service.UserAnalyticServiceIT ‑ test repo get last study history for user()
com.epam.brn.service.UserAnalyticsServiceTest ‑ should not return user with analytics()
com.epam.brn.service.UserAnalyticsServiceTest ‑ should prepareAudioFileMetaData default speed correctly for one word and good statistics()
com.epam.brn.service.UserAnalyticsServiceTest ‑ should prepareAudioFileMetaData normal speed for single word()
com.epam.brn.service.UserAnalyticsServiceTest ‑ should prepareAudioFileMetaData slow correctly for single word and bad statistics()
com.epam.brn.service.UserAnalyticsServiceTest ‑ should prepareAudioFileMetaData with adding comma and slow speed for words with good stat WORDS()
com.epam.brn.service.UserAnalyticsServiceTest ‑ should prepareAudioFileMetaData with adding comma and slowest speed for words with bad stat WORDS ()
com.epam.brn.service.UserAnalyticsServiceTest ‑ should prepareAudioFileMetaData with lera voice up to 18 years old user()
com.epam.brn.service.UserAnalyticsServiceTest ‑ should prepareAudioFileMetaData without adding comma and slow speed for words with good stat PHRASES()
com.epam.brn.service.UserAnalyticsServiceTest ‑ should prepareAudioFileMetaData without adding comma and slowest speed for words with bad stat PHRASES()
…
com.epam.brn.integration.UserDailyAnalyticsJobIT ‑ should be idempotent when run twice on the same day()
com.epam.brn.integration.UserDailyAnalyticsJobIT ‑ should fill a daily analytics snapshot for a user()
com.epam.brn.integration.service.UserDailyAnalyticServiceIT ‑ test repo get last study history for user()
com.epam.brn.job.UserDailyAnalyticsJobTest ‑ should replace today's snapshot and insert fresh rows()
com.epam.brn.job.UserDailyAnalyticsJobTest ‑ should swallow exceptions so the scheduler keeps running()
com.epam.brn.service.UserDailyAnalyticsServiceTest ‑ should not return user with analytics()
com.epam.brn.service.UserDailyAnalyticsServiceTest ‑ should prepareAudioFileMetaData default speed correctly for one word and good statistics()
com.epam.brn.service.UserDailyAnalyticsServiceTest ‑ should prepareAudioFileMetaData normal speed for single word()
com.epam.brn.service.UserDailyAnalyticsServiceTest ‑ should prepareAudioFileMetaData slow correctly for single word and bad statistics()
com.epam.brn.service.UserDailyAnalyticsServiceTest ‑ should prepareAudioFileMetaData with adding comma and slow speed for words with good stat WORDS()
…

♻️ This comment has been updated with latest results.

@andrsam andrsam changed the title 2609 [BE] filling analytics table using job 2609 [BE] new analytics functionality using database table Mar 30, 2025
@andrsam andrsam changed the title 2609 [BE] new analytics functionality using database table 2609 [BE] new user analytics functionality using database table Mar 30, 2025
@andrsam
andrsam force-pushed the 2609 branch 2 times, most recently from 57db0fc to d82ef2f Compare March 30, 2025 10:03
andrsam added 3 commits March 31, 2025 19:20
# Conflicts:
#	src/main/kotlin/com/epam/brn/controller/UserDetailController.kt
#	src/main/kotlin/com/epam/brn/service/impl/UserAnalyticsServiceImpl.kt
@sonarqubecloud

Copy link
Copy Markdown

@ElenaSpb ElenaSpb self-assigned this Aug 26, 2025
@sonarqubecloud

Copy link
Copy Markdown

@ElenaSpb ElenaSpb linked an issue Aug 26, 2025 that may be closed by this pull request
Comment thread src/main/kotlin/com/epam/brn/model/projection/UsersWithAnalyticsView.kt Outdated
Comment thread src/main/kotlin/com/epam/brn/repo/UserAnalyticsRepository.kt Outdated
Comment thread src/main/kotlin/com/epam/brn/service/impl/UserAnalyticsServiceV1Impl.kt Outdated
Comment thread src/main/kotlin/com/epam/brn/service/UserAnalyticsServiceV1.kt Outdated
Comment thread src/main/kotlin/com/epam/brn/config/SwaggerConfig.kt
Comment thread src/main/kotlin/com/epam/brn/config/SwaggerConfig.kt Outdated
@github-actions

Copy link
Copy Markdown

Frontend test coverage: 45.59%

🤷‍♂️ Did not change

…perseded endpoint path

Variant A review outcome for PR #2687:
- Keep the author's core: nightly UserAnalyticsJob + user_analytics table + aggregation SQL.
- Repurpose the table as a daily snapshot history (snapshot_date, idempotent re-run);
  the live /admin/users endpoint stays realtime (N+1 already fixed on master by #2971).
- Actualize to jakarta.* (Spring Boot 3.5) and the V2yearmonthday migration naming.
- Drop the now-superseded serve-from-table path: UserAnalyticsServiceV1(+Impl),
  UsersWithAnalyticsView projection, UserDetailControllerConfig feature flag,
  the comparison IT, and the javax-era migration.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@ElenaSpb
ElenaSpb self-requested a review October 9, 2026 06:41
The per-test @AfterEach only cleared user_analytics and user_account,
leaving the exercise_group/series/subgroup/exercise content behind. The
second test then re-inserted the default series and hit a unique-name
constraint on exercise_group. Use BaseIT.deleteInsertedTestData() to
tear down the full content graph in the correct FK order.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown

Frontend test coverage: 72.72% (-0.02% compared to 72.74% on base)

@ElenaSpb

ElenaSpb commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Direction for the user-analytics precompute: per-day delta + near-real-time freshness

Two design points we'd like to adopt as this moves forward.

  1. Store a per-day delta, not a cumulative snapshot

The current daily aggregation is cumulative — each row holds the all-time total as of that date (the query aggregates the whole study_history with no per-day filter). Proposal: scope each row to the activity of that day only
(WHERE start_time >= current_date AND start_time < current_date + 1), so a row means "what this user did on this day."

Why this is better:

  • Roll-ups become trivial. Weekly / monthly / yearly are just SUM(spent_time), SUM(done_exercises) and COUNT(*) of daily rows over the period. No need for separate weekly/monthly/yearly tables — they're derived from the daily
    table (a view/query is enough; materialize later only if measurements demand it).
  • study_days drops out of the table. "Study days in a month" is simply the count of daily rows in that month, so it stops being a stored column and becomes a derived count.
  • spent_time and done_exercises sum cleanly across days; first_done / last_done are per-day first/last activity and are combined with MIN/MAX over a period (not SUM).

All-time values that the cumulative snapshot used to hold (lifetime first study, total time, etc.) move into a separate running per-user row (user_lifetime_analytics), upserted rather than re-scanned.

  1. Keep the current day fresh via event-driven incremental updates

Instead of only filling analytics nightly, update today's delta row immediately when an exercise is completed:

  • Hook on study_history commit via @TransactionalEventListener(phase = AFTER_COMMIT) + @async — so it never adds latency to the exercise-completion request.

The current daily aggregation is cumulative — each row holds the all-time total as of that date (the query aggregates the whole study_history with no per-day filter). Proposal: scope each row to the activity of that day only
(WHERE start_time >= current_date AND start_time < current_date + 1), so a row means "what this user did on this day."

Why this is better:

  • Roll-ups become trivial. Weekly / monthly / yearly are just SUM(spent_time), SUM(done_exercises) and COUNT(*) of daily rows over the period. No need for separate weekly/monthly/yearly tables — they're derived from the daily
    table (a view/query is enough; materialize later only if measurements demand it).
  • study_days drops out of the table. "Study days in a month" is simply the count of daily rows in that month, so it stops being a stored column and becomes a derived count.
  • spent_time and done_exercises sum cleanly across days; first_done / last_done are per-day first/last activity and are combined with MIN/MAX over a period (not SUM).

All-time values that the cumulative snapshot used to hold (lifetime first study, total time, etc.) move into a separate running per-user row (user_lifetime_analytics), upserted rather than re-scanned.

  1. Keep the current day fresh via event-driven incremental updates

Instead of only filling analytics nightly, update today's delta row immediately when an exercise is completed:

  • Hook on study_history commit via @TransactionalEventListener(phase = AFTER_COMMIT) + @async — so it never adds latency to the exercise-completion request.
  • Apply an atomic increment in the DB, not a read-modify-write, to stay safe under concurrent completions:
    INSERT INTO user_daily_analytics (snapshot_date, user_id, role_name, done_exercises, spent_time, first_done, last_done)
    VALUES (current_date, :uid, :role, 1, :sec, :ts, :ts)
    ON CONFLICT (user_id, role_name, snapshot_date)
    DO UPDATE SET done_exercises = user_daily_analytics.done_exercises + 1,
    spent_time = user_daily_analytics.spent_time + EXCLUDED.spent_time,
    last_done = GREATEST(user_daily_analytics.last_done, EXCLUDED.last_done);

The nightly job then changes role from filler to reconciler: once a night it recomputes each day's delta from the raw history and overwrites it, self-healing any drift from missed or duplicated events. Net result: incremental
updates give intra-day freshness, the nightly recompute guarantees correctness (belt-and-suspenders).

Note this only works cleanly with the delta model — with a cumulative snapshot, an incremental update would mean recomputing all-time aggregates on every exercise. The delta model makes it a simple +1 / + seconds.

Rollout order: delta migration for user_daily_analytics → event-driven incremental listener → user_lifetime_analytics for all-time columns → read views for week/month/year.

@github-actions

Copy link
Copy Markdown

Frontend test coverage: 72.74% (-0.03% compared to 72.77% on base)

@sonarqubecloud

Copy link
Copy Markdown

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BE] Optimize studyHistory aggregation

2 participants