video_manager now pools only the ffmpeg reader and writer. read_video_frame
seeks and reads through the reader; count_video_frame_total / detect_video_fps
/ detect_video_resolution read from ffprobe metadata (cached via lru_cache).
Removes get_video_capture, conditional_set_video_frame_position, the cv2 import
and the 'capture' pool key (and VideoCaptureSet). Inline the reader buffer
margin as a local (buffer_margin = 16) instead of a module constant.
Side effect: reference face-selector mode and the NSFW analyser now decode AV1
correctly, since they read frames through the reader instead of cv2.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two selectable workflows share one swap loop (process_target_frame):
- disk (default): extract -> process -> merge; temp frames on disk, supports
fps resampling, robust/inspectable.
- stream: read -> swap -> encode over pipes; no temp frames, ~2x faster wall.
Remove --keep-temp entirely (arg, config, state, ui, locale). Replace the dead
cv2.VideoWriter in video_manager with an ffmpeg rawvideo encoder held in the
pool (symmetric to the reader): get_video_writer / write_video_writer_frame /
close_video_writer, with fix_video_encoder relocated to ffmpeg_builder.
The encoder must be closed before processor post_process(), which clears the
video pool and would otherwise SIGTERM the still-open writer.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Instead of writing each swapped frame to a temp PNG and re-encoding them in a
separate merge pass, open the output encoder up front and feed swapped frames
straight into it over a pipe. process_video keeps a bounded look-ahead
(execution_thread_count * 2) and consumes futures oldest-first, so frames reach
the encoder in order with parallelism preserved and memory bounded.
Removes the temp PNG write (process) and the whole merge phase. Wall time
~19 -> ~15.5 s AV1 and ~22 -> ~18 s H.264 for 300 frames -- about half of
3.6.1 -- output unchanged. The old extract_frames/merge_frames stay defined
but unused.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
With the canvas coming from the video reader, the extracted temp PNGs were
never read for pixels -- extraction only produced write targets and a frame
list. Drop the extract_frames step entirely: process_video iterates the trim
range, each worker reads its frame from the reader, swaps, and writes straight
to the temp frame path that merge consumes.
Removes the whole extracting phase (a second full-video decode + PNG encode).
Wall time drops ~25-36% (AV1 ~30 -> ~19 s, H.264 ~29 -> ~22 s for 300 frames);
processing throughput unchanged; output unchanged.
Note: this path does not resample fps (output_video_fps must equal source fps)
and derives the frame count via cv2 container metadata; a scaled output falls
back to a per-frame resize instead of the temp PNG.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
process_temp_frame read the target frame twice per output frame: once from
the video (the detection window) and once from the extracted temp PNG (the
swap canvas). They are the same frame N. The window centre is already decoded
in the reader, so use it as the canvas and only fall back to the temp PNG when
the reader resolution differs from the temp resolution (scaled output).
Removes one full-frame decode per output frame. AV1 ~22 -> ~26 fps and
H.264 ~20 -> ~23 fps -- matching and slightly beating 3.6.1, because a
frame copy is cheaper than a PNG decode. The canvas now comes from the raw
pipe decode instead of the PNG decode of the same frame (sub-visual ~1 LSB
difference in passthrough regions); detection is unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
read_video_reader_frame copied every decoded frame out of the pipe buffer.
The frames are only read downstream (face detection, coordinate scaling), so
the read-only numpy view over the pipe bytes is safe to return directly.
Saves one full-frame (~25 MB at 4K) memcpy per decoded frame.
AV1 ~21 -> ~22 fps, H.264 ~19 -> ~20 fps, output unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
select_video_frames routed the target window through read_static_video_chunk
(lru_cache of chunk-aligned 5-frame blocks), decoding ~2 chunks / 10 frames
to serve a 5-frame window. With the persistent sequential ffmpeg reader this
indirection is unnecessary.
Keep a small sliding frame buffer on the reader instance and read the window
straight from it: sequential processing decodes ~1 new frame per output frame
(matching 3.6.1), no chunk-alignment waste. A margin absorbs the thread
out-of-order access without pipe restarts; a backward or large-forward miss
restarts the pipe at the target frame. The old chunk functions stay in place
(now unused by the window path) so processor cache_clear calls keep working.
AV1 ~20 -> ~21 fps, H.264 ~18 -> ~19 fps, and lower RAM (~12 -> ~9.7 GB),
output pixel-identical to 3.6.1.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
video_manager built its video reader on cv2.VideoCapture, which cannot
decode AV1 with the bundled opencv wheel on Linux, so the multi-window
target pack (read_video_chunk -> select_video_frames) came back empty and
frames passed through unswapped.
Replace the reader with a persistent ffmpeg rawvideo pipe whose instance
lives in the pool store, mirroring the cv2.VideoCapture lifecycle: probe
width/height/fps/frame_total via ffprobe, spawn an ffmpeg pipe seeked to
the chunk start, read frames sequentially, restart on a backward seek.
All ffmpeg/ffprobe command construction goes through the builders; the
single-frame cv2 path is left untouched (scope is the multi-window reader).
ffprobe.py / ffprobe_builder.py follow the v4 template. AV1 now swaps
correctly at ~20 fps (vs the ~11 fps temp-frame variant); H.264 stays
correct at ~18 fps.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* patch 3.7.1
* fix bug for 2 processors on image to image pipeline (#1177)
* reduce default value of target_frame_amount
* restore old performance
* restore old performance
* restore old performance
* restore old performance
* mark as next, introduce dynamic scale for face debugger
* use latest onnxruntime
* update within Gradio 5
* Remove system memory limit (#986)
* remove system memory limit from ui
* remove system memory limit from args.py
* flatten the face store
* prevent countless importlib.import_module calls
* remove --onnxruntime from install.py
* remove --onnxruntime from install.py
* resolve static inference providers to fix macos (#1127)
* resolve static inference providers to fix macos
* fix lint
* restore old behaviour
* restore old behaviour
* handle ghost and uniface as well
* adjust condition for ghost and uniface
* fix Gradio gallery styles
* remove face store (#1132)
* fix dataflow in streamer
* Face selector auto mode (#1137)
* introduce face selector auto mode
* introduce face selector auto mode
* introduce face selector auto mode
* correct way is to pass source_vision_frames
* make the world a better place
* fix dataflow in faceswapper, no read of files withing inner methods (#1148)
* fix dataflow in faceswapper, no read of files withing inner methods
* fix lint
* adjust code more
* adjust code more
* bring back the face store but for source and reference only (#1149)
* bring back the face store but for source and reference only
* fix ci
* minor improvement
* guard for tobytes()
* drop condition in select_faces()
* Replace CONFIG_PARSER global with @lru_cache (#1147)
* remove global config_parser
* fix import order
* remove lambda
* remove unused block
* optimize app context detection
* decouple common modules from core (#1152)
* decouple common modules from core
* remove that nonsense
* remove that nonsense
* minor adjustment to workflows
* Tag HEVC output as hvc1 and move moov atom to the front (#1153)
* Tag HEVC output as hvc1 and move moov atom to the front
ffmpeg defaults HEVC in MP4 to the 'hev1' sample entry and leaves the moov
atom at the tail. Apple players (QuickTime, Finder QuickLook) refuse to decode
'hev1' and stall reading a tail-placed moov on large files, so hevc_nvenc /
libx265 renders cannot be previewed on macOS.
- add ffmpeg_builder.set_video_tag(): emit `-tag:v hvc1` for every HEVC
encoder (libx265, hevc_nvenc, hevc_amf, hevc_qsv, hevc_videotoolbox).
Applied in merge_video where the encoder is known; `-c:v copy` in the audio
mux / concat steps preserves the tag.
- add ffmpeg_builder.set_faststart(): emit `-movflags +faststart`, applied in
restore_audio / replace_audio / concat_video which write the final output.
H.264 and other codecs are left untouched. Verified on a real hevc_nvenc
render: hev1 hung QuickLook (no thumbnail); after the patch the file is hvc1
with a front-placed moov and QuickLook generates a thumbnail.
* Restrict hvc1 tag and faststart to quicktime containers
Gate set_video_tag / set_faststart on the output container format
(m4v, mov, mp4) via get_file_format(), so non-quicktime muxers no longer
receive -tag:v hvc1 / -movflags +faststart. Trim test_set_video_tag to a
single positive and negative assertion.
Addresses review on #1153.
* Move hvc1 tag and faststart gates into ffmpeg_builder
Rename set_video_tag / set_faststart to conditional_* and push the
container-format gate (m4v, mov, mp4) inside the builders, keeping
ffmpeg.py free of inline conditionals. Matches the set_image_quality
pattern. Addresses review on #1153.
* post cleanup after merge
* Pack target frames (#1158)
* pack target frames
* add todos
* add todos, resolve todos
* resolve todos
* change names
* revert to single target frame for select faces
* fix lint
* return empty frame
* get() have no default
* Fix trim (#1162)
* fix trim
* fix trim
* rename ffmpeg builder method
* rename to temp_frame_set and temp_frame_pattern
---------
Co-authored-by: harisreedhar <h4harisreedhar.s.s@gmail.com>
Co-authored-by: Harisreedhar <46858047+harisreedhar@users.noreply.github.com>
* Implement face tracker (#1163)
* add face tracker
* change get_nearest_track_face -> get_nearest_track_index
* create face_creator.py and move methods around
* add type FaceTrack
* naming
* remove iou test, don't belong there
* fix spaces
* rename to interpolate_points
* rename to find_best_face_track
* just track_faces
* cleanp
* previous next naming
* remove >= and >=
* rename
* remove helper from test and use face from source.jpg
* make get_anchor_indices more readable
* track_faces() call before and is forwarded to select_faces
* change to interpolate_faces
* rename methods
* rename methods
* rename variables
* remove dtype
* move face_anlyser -> face_creator
* claenup face_creator.py
* move tests to dedicated test face detector
* move tracking inside select_faces
* simplify face_tracker (#1165)
* minor renaming
* improve face_tracker test (#1166)
* improve face_tracker test
* cleanup
* Add target frame amount (#1167)
* introduce --target-frame-amount
* add ui
* make track_faces conditional
* update choices.py
* fix []
* rename component file to frame_process.py
* fix track preview (#1168)
* introduce face origin (#1169)
* add guard to prevent failure
* show and hide voice extractor according to lip syncer
* rename average_face_coordinates to average_face_geometry
* use static faces for select_faces()
* face store with lock (#1171)
* face store with lock
* face store with lock
* remove refill color from bbox
* adjust tests and handle frame_position proper way
* enforce similar naming
* introduce face tracker score
* introduce face tracker score
* fix/audio-trim-alignment (#1173)
* fix audio offset
* fix audio offset
* remove reference_frame_number check
---------
Co-authored-by: harisreedhar <h4harisreedhar.s.s@gmail.com>
* reduce face tracker score from 0 to 0.5
* mark as 3.7.0
* make face tracker stateless
---------
Co-authored-by: Harisreedhar <46858047+harisreedhar@users.noreply.github.com>
Co-authored-by: kazuki nakai <kazuki.nakai@agiletec.net>
Co-authored-by: harisreedhar <h4harisreedhar.s.s@gmail.com>
* Avast intercepts ssl cert and breaks curl
* modernize dependencies
* fix sanitizer for int range
* comment out tests
* disable testing for CI
* disable testing for CI
* mark as next
* add fran model
* add support for corridor key (#1060)
* introduce despill color
* simplify the apply dispill color
* finalize naming for both fill and despill
* follow vision_frame convension
* patch fran model
* adjust fran urls
* Feat/dynamic env setup (#1061)
* dynamic environment setup
* dynamic environment setup
* fix fran model
* prevent directml using incompatible corridor_key model
* fix environment setup for windows
* switch to corridor_key_1024 and corridor_key_2048
* switch to corridor_key_1024 and corridor_key_2048
* mark it as 3.6.0
* rename environment to conda
* rename environment to conda
* fix testing for face analyser
* some background remove cosmetics
* some background remove cosmetics
* some background remove cosmetics
* update preview
---------
Co-authored-by: harisreedhar <h4harisreedhar.s.s@gmail.com>
* remove insecure flag from curl
* eleminate repating definitons
* limit processors and ui layouts by choices
* follow couple of v4 standards
* use more secure mkstemp
* dynamic cache path for execution providers
* fix benchmarker, prevent path traveling via job-id
* fix order in execution provider choices
* resort by prioroty
* introduce support for QNN
* close file description for Windows to stop crying
* prevent ConnectionResetError under windows
* needed for nested .caches directory as onnxruntime does not create it
* different approach to silent asyncio
* update dependencies
* simplify the name to just inference providers
* switch to trt_builder_optimization_level 4
Enforce consistent space inside square brackets for all list literals
across source and test files.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude <noreply@anthropic.com>
* honor webcam resolution to avoid stripe mismatch, update dependencies
* avoid version conflicts
* enforce prores video extraction to 8 bit
* make the installer more robust on execution switch
* make the installer more robust on execution switch
* improve the installer env handling
* different approach to handle env
* Mark as NEXT
* Reduce caching to avoid RAM explosion
* Reduce caching to avoid RAM explosion
* Update dependencies
* add face-detector-pad-factor
* update facefusion.ini
* fix test
* change pad to margin
* fix order
* add prepare margin
* use 50% max margin
* Minor fixes part2
* Minor fixes part3
* Minor fixes part4
* Minor fixes part1
* Downgrade onnxruntime as of BiRefNet broken on CPU
add test
update
update facefusion.ini
add birefnet
* rename models
add more models
* Fix versions
* Add .claude to gitignore
* add normalize color
add 4 channel
add colors
* worflows
* cleanup
* cleanup
* cleanup
* cleanup
* add more models (#961)
* Fix naming
* changes
* Fix style and mock Gradio
* Fix style and mock Gradio
* Fix style and mock Gradio
* apply clamp
* remove clamp
* Add normalizer test
* Introduce sanitizer for the rescue (#963)
* Introduce sanitizer for the rescue
* Introduce sanitizer for the rescue
* Introduce sanitizer for the rescue
* prepare ffmpeg for alpha support
* Some cleanup
* Some cleanup
* Fix CI
* List as TypeAlias is not allowed (#967)
* List as TypeAlias is not allowed
* List as TypeAlias is not allowed
* List as TypeAlias is not allowed
* List as TypeAlias is not allowed
* Add mpeg and mxf support (#968)
* Add mpeg support
* Add mxf support
* Adjust fix_xxx_encoder for the new formats
* Extend output pattern for batch-run (#969)
* Extend output pattern for batch-run
* Add {target_extension} to allowed mixed files
* Catch invalid output pattern keys
* alpha support
* cleanup
* cleanup
* add ProcessorOutputs type
* fix preview and streamer, support alpha for background_remover
* Refactor/open close processors (#972)
* Introduce open/close processors
* Add locales for translator
* Introduce __autoload__ for translator
* More cleanup
* Fix import issues
* Resolve the scope situation for locals
* Fix installer by not using translator
* Fixes after merge
* Fixes after merge
* Fix translator keys in ui
* Use LOCALS in installer
* Update and partial fix DirectML
* Use latest onnxruntime
* Fix performance
* Fix lint issues
* fix mask
* fix lint
* fix lint
* Remove default from translator.get()
* remove 'framerate='
* fix test
* Rename and reorder models
* Align naming
* add alpha preview
* fix frame-by-frame
* Add alpha effect via css
* preview support alpha channel
* fix preview modes
* Use official assets repositories
* Add support for u2net_cloth
* fix naming
* Add more models
* Add vendor, license and year direct to the models
* Add vendor, license and year direct to the models
* Update dependencies, Minor CSS adjustment
* Ready for 3.5.0
* Fix naming
* Update about messages
* Fix return
* Use groups to show/hide
* Update preview
* Conditional merge mask
* Conditional merge mask
* Fix import order
---------
Co-authored-by: harisreedhar <h4harisreedhar.s.s@gmail.com>
Co-authored-by: Harisreedhar <46858047+harisreedhar@users.noreply.github.com>
* Fix preview when using frame enhancer
* Fix version conflict numpy vs. cv2
* Use latest numpy
* Introduce scale_face() to match size of temp frames and target frames
* Remove hardcoded backend for camera under Windows
* Up and downgrade some dependencies
* Up and downgrade some dependencies
* Up and downgrade some dependencies
* Rename calcXXX to calculateXXX
* Add migraphx support
* Add migraphx support
* Add migraphx support
* Add migraphx support
* Add migraphx support
* Add migraphx support
* Use True for the flags
* Add migraphx support
* add face-swapper-weight
* add face-swapper-weight to facefusion.ini
* changes
* change choice
* Fix typing for xxxWeight
* Feat/log inference session (#906)
* Log inference session, Introduce time helper
* Log inference session, Introduce time helper
* Log inference session, Introduce time helper
* Log inference session, Introduce time helper
* Mark as NEXT
* Follow industry standard x1, x2, y1 and y2
* Follow industry standard x1, x2, y1 and y2
* Follow industry standard in terms of naming (#908)
* Follow industry standard in terms of naming
* Improve xxx_embedding naming
* Fix norm vs. norms
* Reduce timeout to 5
* Sort out voice_extractor once again
* changes
* Introduce many to the occlusion mask (#910)
* Introduce many to the occlusion mask
* Then we use minimum
* Add support for wmv
* Run platform tests before has_execution_provider (#911)
* Add support for wmv
* Introduce benchmark mode (#912)
* Honestly makes no difference to me
* Honestly makes no difference to me
* Fix wording
* Bring back YuNet (#922)
* Reintroduce YuNet without cv2 dependency
* Fix variable naming
* Avoid RGB to YUV colorshift using libx264rgb
* Avoid RGB to YUV colorshift using libx264rgb
* Make libx264 the default again
* Make libx264 the default again
* Fix types in ffmpeg builder
* Fix quality stuff in ffmpeg builder
* Fix quality stuff in ffmpeg builder
* Add libx264rgb to test
* Revamp Processors (#923)
* Introduce new concept of pure target frames
* Radical refactoring of process flow
* Introduce new concept of pure target frames
* Fix webcam
* Minor improvements
* Minor improvements
* Use deque for video processing
* Use deque for video processing
* Extend the video manager
* Polish deque
* Polish deque
* Deque is not even used
* Improve speed with multiple futures
* Fix temp frame mutation and
* Fix RAM usage
* Remove old types and manage method
* Remove execution_queue_count
* Use init_state for benchmarker to avoid issues
* add voice extractor option
* Change the order of voice extractor in code
* Use official download urls
* Use official download urls
* add gui
* fix preview
* Add remote updates for voice extractor
* fix crash on headless-run
* update test_job_helper.py
* Fix it for good
* Remove pointless method
* Fix types and unused imports
* Revamp reference (#925)
* Initial revamp of face references
* Initial revamp of face references
* Initial revamp of face references
* Terminate find_similar_faces
* Improve find mutant faces
* Improve find mutant faces
* Move sort where it belongs
* Forward reference vision frame
* Forward reference vision frame also in preview
* Fix reference selection
* Use static video frame
* Fix CI
* Remove reference type from frame processors
* Improve some naming
* Fix types and unused imports
* Fix find mutant faces
* Fix find mutant faces
* Fix imports
* Correct naming
* Correct naming
* simplify pad
* Improve webcam performance on highres
* Camera manager (#932)
* Introduce webcam manager
* Fix order
* Rename to camera manager, improve video manager
* Fix CI
* Remove optional
* Fix naming in webcam options
* Avoid using temp faces (#933)
* output video scale
* Fix imports
* output image scale
* upscale fix (not limiter)
* add unit test scale_resolution & remove unused methods
* fix and add test
* fix
* change pack_resolution
* fix tests
* Simplify output scale testing
* Fix benchmark UI
* Fix benchmark UI
* Update dependencies
* Introduce REAL multi gpu support using multi dimensional inference pool (#935)
* Introduce REAL multi gpu support using multi dimensional inference pool
* Remove the MULTI:GPU flag
* Restore "processing stop"
* Restore "processing stop"
* Remove old templates
* Go fill in with caching
* add expression restorer areas
* re-arrange
* rename method
* Fix stop for extract frames and merge video
* Replace arcface_converter models with latest crossface models
* Replace arcface_converter models with latest crossface models
* Move module logs to debug mode
* Refactor/streamer (#938)
* Introduce webcam manager
* Fix order
* Rename to camera manager, improve video manager
* Fix CI
* Fix naming in webcam options
* Move logic over to streamer
* Fix streamer, improve webcam experience
* Improve webcam experience
* Revert method
* Revert method
* Improve webcam again
* Use release on capture instead
* Only forward valid frames
* Fix resolution logging
* Add AVIF support
* Add AVIF support
* Limit avif to unix systems
* Drop avif
* Drop avif
* Drop avif
* Default to Documents in the UI if output path is not set
* Update wording.py (#939)
"succeed" is grammatically incorrect in the given context. To succeed is the infinitive form of the verb. Correct would be either "succeeded" or alternatively a form involving the noun "success".
* Fix more grammar issue
* Fix more grammar issue
* Sort out caching
* Move webcam choices back to UI
* Move preview options to own file (#940)
* Fix Migraphx execution provider
* Fix benchmark
* Reuse blend frame method
* Fix CI
* Fix CI
* Fix CI
* Hotfix missing check in face debugger, Enable logger for preview
* Fix reference selection (#942)
* Fix reference selection
* Fix reference selection
* Fix reference selection
* Fix reference selection
* Side by side preview (#941)
* Initial side by side preview
* More work on preview, remove UI only stuff from vision.py
* Improve more
* Use fit frame
* Add different fit methods for vision
* Improve preview part2
* Improve preview part3
* Improve preview part4
* Remove none as choice
* Remove useless methods
* Fix CI
* Fix naming
* use 1024 as preview resolution default
* Fix fit_cover_frame
* Uniform fit_xxx_frame methods
* Add back disabled logger
* Use ui choices alias
* Extract select face logic from processors (#943)
* Extract select face logic from processors to use it for face by face in preview
* Fix order
* Remove old code
* Merge methods
* Refactor face debugger (#944)
* Refactor huge method of face debugger
* Remove text metrics from face debugger
* Remove useless copy of temp frame
* Resort methods
* Fix spacing
* Remove old method
* Fix hard exit to work without signals
* Prevent upscaling for face-by-face
* Switch to version
* Improve exiting
---------
Co-authored-by: harisreedhar <h4harisreedhar.s.s@gmail.com>
Co-authored-by: Harisreedhar <46858047+harisreedhar@users.noreply.github.com>
Co-authored-by: Rafael Tappe Maestro <rafael@tappemaestro.com>